CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project Overview
Strum VOD is a self-hosted VOD (Video on Demand) platform. It handles video upload (standard presigned-URL or resumable TUS), transcoding to a multi-quality HLS ladder (360p–4320p), and playback delivery via Cloudflare R2 (both local dev and production).
Package Manager
This project uses pnpm with workspaces. Always use pnpm commands, not npm.
Commands
Cloudflare (wrangler — primary deploy path)
# First-time setup
cp .env.example .env
pnpm install
# R2 buckets + CORS (see infra/r2/)
pnpm r2:list # list buckets
pnpm r2:cors:list # show CORS rules on the strum-vod bucket
pnpm r2:cors:set # apply infra/r2/strum-vod-cors.json
pnpm r2:cors:delete # clear CORS rules
# Frontends
pnpm deploy:all # build + deploy EVERYTHING (api-edge, tus-edge, dashboard, player, landing, docs, worker, transcoder)
pnpm deploy:all cloudflare # Cloudflare only (api-edge, tus-edge, dashboard Worker `video-on-demand`, player/landing/docs Pages)
pnpm deploy:all fly # Fly only (worker, transcoder)
# Individual deploys
pnpm --filter @strum-vod/api-edge run deploy # api-edge → Cloudflare Worker `strum-vod-api-edge` (serves vapi.usestrum.app — THE backend)
pnpm --filter @strum-vod/tus-edge run deploy # tus-edge → Cloudflare Worker (resumable TUS upload)
wrangler deploy --config apps/dashboard/wrangler.toml # dashboard → Cloudflare Worker `video-on-demand` (serves vod.usestrum.app)
pnpm deploy:player # player app (wrangler pages → Cloudflare Pages project `player`)
pnpm deploy:landing # landing (wrangler pages → Cloudflare Pages project `strum-vod-landing`)
pnpm deploy:docs # docs (wrangler pages → Cloudflare Pages project `strum-vod-docs`)
fly deploy --config fly.worker.toml # Worker (Fly `strum-vod-worker`)
fly deploy --config fly.transcoder.toml # Transcoder (Fly `strum-vod-transcoder`)
# A worker binary running anywhere else (operator's own VPS/laptop) isn't
# deployed by any of the above — see docs/worker-setup.md.Per-app development (requires local PostgreSQL, Redis, and R2 credentials in .env)
pnpm run build # build all workspaces
pnpm run typecheck # typecheck all workspaces
# Start everything at once (uses concurrently, color-coded output per process)
pnpm run dev
# Or start subsets:
pnpm run dev:web # dashboard + player only (frontend-only work)
pnpm run dev:api # API only
pnpm run dev:transcoder # Transcoder only (Go: ffmpeg ladder — needs local Go toolchain)
pnpm run dev:dashboard # Dashboard only
pnpm run dev:player # Player app only
pnpm run dev:qstash # Local QStash emulator (@upstash/qstash-cli dev, no tunnel/account needed) — see ".env.example"'s "QStash local dev" block
# Individual builds
pnpm run build -w @strum-vod/db # must run before worker
pnpm run build -w @strum-vod/dashboard
pnpm run build:player
pnpm run build:transcoder # apps/transcoder (Go) — needs Go 1.23+ toolchain, not a pnpm workspaceDocs site (VitePress)
pnpm docs:dev # dev server with hot reload (default: http://localhost:5173)
pnpm docs:build # static build → docs/.vitepress/dist
pnpm docs:preview # serve the built site (default: http://localhost:4173)
pnpm deploy:docs # build + wrangler pages deploy → Cloudflare Pages project strum-vod-docsTesting
apps/api (Fly, Fastify) is deleted (2026-08-16) along with its Testcontainers E2E/unit suites — apps/api-edge (the live backend, Cloudflare Workers) has no automated test suite of its own beyond typecheck; the closest thing to integration coverage today is pnpm test:e2e:edge (see below), a shallower end-to-end smoke test through the real edge stack rather than a fine-grained per-route Testcontainers suite. Porting that coverage to api-edge is a known, real gap — not yet done.
pnpm --filter @strum-vod/cli run test # CLI unit tests (pure logic — no Docker, no DB)
pnpm --filter @strum-vod/ai-pipeline run test # provider unit tests (shared with trigger.dev tasks)
pnpm test:transcoder # Go constant parity testFully removed (2026-08-16). The customer-facing "Nodes" (self-hosted BYO-compute) and admin "Platform Workers" fleets — and every trace of them (
routes/nodes.ts,routes/node-agent.ts,routes/platform-workers.ts,services/nodes.ts,requireNodeAuth(), thenk_/pw_token families,cmd/node-agent(Go), the/admin/nodes//admin/platform-workersdashboard pages,apps/api's dedicatedtest:e2e:node-agentE2E suite, and thenodes/node_jobs/platform_workers/platform_worker_tokenstables plus theorganizations.node_routing_enabled/settings.nodes_enabled/settings.platform_worker_*columns — deleted from bothapps/api-edgeandapps/api(Fly), schema included. All superseded by the unifiedmachines/machineTokens/dispatchJobsfleet (/admin/machines, see "Worker unification" below) — one binary (cmd/transcoder), one broker (/v1/worker-agent/*), one credential type (mt_live_...), for every machine regardless of whether it's the Fly fleet or an operator's own VPS.
Live scale-to-zero smoke test (scripts/smoke-test-scale-to-zero.sh, pnpm smoke:scale-to-zero) — the one test that isn't Testcontainers-based by design: it hits the deployed strum-vod-api/strum-vod-transcoder Fly apps for real, because scale-to-zero is a Fly Machines runtime behavior (proxy autostart, internal/selfstop's self-exit) that no local container can reproduce. Requires fly, curl, jq on PATH and API_KEY=<org API key> (X-Api-Key). It POSTs /v1/assets/demo-seed (real public sample video, seedFakeAi: true so it needs no AI provider keys), then polls fly machine list --app <app> --json through the full cycle: transcoder stopped→started (API's prewarmFleet wake call) → asset ready (real ffmpeg encode) → transcoder started→stopped again (internal/selfstop's idle self-exit, ~SELF_STOP_IDLE_SECONDS later). Tails the transcoder's fly logs in the background and prints the [selfstop]/[*-wake] lines at the end regardless of outcome, plus phase timings (process → transcoder started, process → ready) to quantify cold-start regressions. Deletes the seeded asset on exit (KEEP_ASSET=1 to keep it for manual inspection). Timeouts are env-overridable (WAKE_TIMEOUT_SECONDS, TRANSCODE_TIMEOUT_SECONDS, SHUTDOWN_TIMEOUT_SECONDS) — see the script's header comment for the full list.
Scheduled DB backups (trigger.dev scheduled task, db-backup) — the daily backup is a trigger.dev scheduled task (cron 0 5 * * *), replacing the retired QStash Schedule and Fly Scheduled Machine backup-cron. The trigger.dev task runs runDbBackup('schedule') + cleanupOldBackups(30) directly (same logic as before); POST /v1/backups/run (manual trigger) calls triggerTask(c.env, 'db-backup-manual') via the REST API. The trigger.dev image installs postgresql-client-18 via a PGDG build layer — the Neon server is Postgres 18, and Debian bookworm's stock pg_dump 15 fatally refuses to dump a newer server ("server version mismatch").
Orchestration migration — LIVE in production (apps/trigger-tasks) — a workspace package holding trigger.dev tasks that replace the QStash-jobType-branching + atomic-claim reaper pattern described above. apps/api-edge's dispatch code unconditionally calls these tasks (no rollout flag — it was removed). Deployed for real: trigger.dev project proj_nzwqjxttfrlgsrahqvyv ("VOD", org strum-latam-32cc) has its prod environment fully provisioned (DATABASE_URL, S3_*, DEEPGRAM_API_KEY, API_PUBLIC_URL, TRANSCRIPTION_PROVIDER — see trigger.config.ts's syncEnvVars extension, which is the only way to set them since the CLI has no env set, only list/get/pull), TRIGGER_SECRET_KEY is set on the strum-vod-api-edge Cloudflare Worker, and a real trigger.dev deploy has run successfully (9 tasks). Verified end to end against real production infra — see the "Worker unification" section below for the full smoke-test result.
The AI pipeline is broken into one task per step, not one monolithic task — each is independently re-triggerable and independently retried by trigger.dev's own attempt mechanism, and each resolves its own aiConfig/transcript/aiOptions from the DB (payload is just {assetId, aiJobId}, plus targetLang for translation) rather than receiving them from a parent:
ai-pipeline— orchestrator only. Runstranscribe-audioviatriggerAndWait(everything else needs its output), then the four mutually-independent post-transcription steps viabatch.triggerByTaskAndWait(a static fan-out of different task types — notbatchTriggerAndWait, which is for the same task with different payloads), thentranslate-subtitlesonce per target language viabatchTriggerAndWait(parallel, unlike the old single-call loop that translated every language sequentially).transcribe-audio— the sync/async provider chain + retry loop +wait.forTokenfor async (modal) dispatch.generate-subtitles,generate-chapters,generate-smart-metadata,discover-highlights— each readsai_jobs.transcript_databack out of the DB itself.discover-highlights— only the LLM discovery half (transcript in, candidate moments out, no ffmpeg). The other half — cutting the clips — is not a trigger.dev task: it dispatches aRENDER_HIGHLIGHTSjob onto the unified worker fleet (dispatch_jobs/routes/worker-agent.ts, same mechanismtranscode-orchestrationuses) and waits on the resulting token. The old/internal/ai/generate-highlightsHTTP endpoint (on the deleted apps/worker) has no caller left — analytics/backup/org-delete are now trigger.dev scheduled tasks. Long-video chunking (Scout/Curator): videos ≥LONG_VIDEO_THRESHOLD_SEC(1800s/30min) runrunScoutCuratorPass()(indiscover-highlights.tsitself) instead of the single-passllm.generateHighlights(). Scout and Curator are two independent, single-shotLlmProvidermethods —generateScoutCandidates(transcriptChunk, chunkDurationSec, resolvedPrompt?)andcurateHighlightCandidates(candidateBlocksText, candidateCount, resolvedPrompt?)— each with its ownprompts.define()id (generate-highlights-scout/generate-highlights-curatorinapps/trigger-tasks/src/prompts/highlights.ts, dashboard-overridable and independently resolvable, same asgenerate-highlights). Config parity with the short-video path:highlightsScoutPromptcarries the sameuserInstruction/contentType/audioFeaturesvariables (and the same content-type-criteria Handlebars blocks) ashighlightsPrompt—runScoutCuratorPass()takes anoptionsparam and threads them through, so an asset's configured content type, custom instruction, and audio-feature analysis (skipped under thefastpreset, same as the short-video path) still apply on a long video. Curator never receives them, on either path — it only judges already-found candidates. Content-type auto-detect:contentTypealmost never gets set on a fully-automatic upload (the dashboard's content-type dropdown,AiOptionsForm.tsx, is opt-in and defaults empty), which left the content-type-criteria prompt blocks dead weight for most real assets.discover-highlights.tsnow callsllm.detectContentType()(one cheap classification call,gpt-4o-mini, its owndetect-highlights-content-typeprompt) over a short sample —buildContentTypeSample()inhighlights.ts, first 25 transcript segments capped at 3000 chars, not the full transcript — whenever the asset has no manually-set contentType; a manual selection is never overridden. Mirrors AI-Youtube-Shorts-Generator'sdetect_content_type(). Runs once, before the short/long branch decision, so both paths benefit. Audio features are no longer partly fabricated:analyzeAudioFeatures()(packages/ai-pipeline/src/audio-analysis.ts) used to also returnlaughterSegmentsandprosodyFeatures.{avgPitch,pitchVariance,speakingRate}— none of that was real signal.laughterSegmentswassilenceSegments's ownsilencedetectffmpeg filter run a second time with a shorter duration window and atype: 'laughter'label slapped on it (no laughter classifier involved); the three prosody fields were hardcoded literals (150/50/3.5) returned unconditionally regardless of the actual audio. Both were removed (2026-08-24) rather than fixed — implementing real laughter/pitch detection was out of scope, and shipping fabricated numbers to the LLM as if they were measured is worse than not sending them. The survivingAudioFeaturesshape isrmsProfile,silenceSegments,spectralFluxPeaks(all genuinely computed from ffmpeg output) pluspeakLevelVarianceDb(the one real fieldprosodyFeaturesever had, promoted to the top level).formatAudioFeaturesForPrompt()and both highlight prompts' "Use audio features to boost scores" rule were updated to match — they no longer promise a laughter signal that never existed. Neither method chunks, formats, or merges anything itself — all of that is pure, LLM-free, and lives inpackages/ai-pipeline/src/highlights.ts(unit-tested inhighlights.test.ts, no mocking required):chunkTranscriptForScout(transcript, durationSec)— splits into ≤SCOUT_MAX_CHUNK_CHARS(18000) pieces, each with a proportional time offset; a transcript under the limit is one whole-video chunk.offsetScoutCandidates(candidates, offsetSec)— shifts a chunk's Scout candidates back onto the full timeline.buildCandidateBlocks(candidates, transcript, durationSec)— formats Scout candidates (each with its own timestamp-sliced snippet) into the text block Curator reads.mergeCuratorDecisions(decisions, candidates)— resolves keptCuratorDecisions back onto real timestamps viacandidateIndices(never array position — Curator drops candidates, so its output array is shorter than and re-indexed relative to Scout's); a decision referencing multiple indices spans from the earliest start to the latest end.
runScoutCuratorPass()is the only piece that wires the two LLM calls and four pure helpers together (one Scout call per chunk, then one Curator call over all chunks' merged candidates) — mirrors AI-Youtube-Shorts-Generator's long-video chunking. Below the threshold, the existing single-passgenerateHighlights()(whole transcript, one LLM call, its owngenerate-highlightsprompt) is unchanged. Overlap dedupe:clampHighlightSuggestions(packages/ai-pipeline/src/highlights.ts) — shared by both paths — runsdedupeOverlappingHighlights()first: for candidates overlapping >50% of the shorter one's own duration, only the higher composite-score one (average of the 5scoresdimensions) survives, before bounds-clamping and capping atHIGHLIGHT_MAX_COUNT. Also ported from the same reference project.translate-subtitles— one language per task run; regenerates the source VTT fromai_jobs.transcript_datarather than depending ongenerate-subtitleshaving already run, and writes only its own language intoai_jobs.subtitle_translationsvia an atomicjsonb ||merge (safe against N parallel writers) — the orchestrator owns the aggregatesubtitleTranslationStatusand the consolidatedtranslations-index.jsonafter the batch resolves.separate-audio-stems— unchanged, fire-and-forget, dispatches Modal directly.
packages/ai-pipeline (transcription/LLM provider clients + the clampHighlightSuggestions validator, shared with the Go-side highlight cutter's input contract) remains the single source of truth for provider logic — no @trigger.dev/sdk outside apps/trigger-tasks (raw REST calls only from apps/api-edge).
Per-step LLM model overrides (2026-08-24): each of the 7 LLM-calling steps (generateChapters, generateHighlights, generateScoutCandidates, curateHighlightCandidates, detectContentType, generateSmartMetadata, translateSegment — all in packages/ai-pipeline/src/providers/ai-sdk-llm.ts) can run on a different model than the shared global llmModel, via the same env+DB-override pattern as deepgramModel/whisperModel/highlightsPreset: AiProviderSettings (packages/db/src/constants.ts) gains chaptersModel/highlightsModel/contentTypeDetectionModel/smartMetadataModel/translationModel, merged in resolveAiConfig() the same dbConfig?.field || envDefaults.field way, surfaced in Admin AI Settings (AdminAiSettingsPage.tsx's new "Modelos por Etapa" block) and via env vars (LLM_CHAPTERS_MODEL etc, read in both apps/trigger-tasks/src/lib/ai-config.ts and apps/api-edge/src/services/ai.service.ts's own getEnvDefaults — two separate files with the same env-reading gap, both fixed). highlights is shared by all three highlight-generation methods (single-pass, Scout, Curator — same task at different video lengths, not three separate knobs).
This replaced a real bug, not just added a feature: every generateObject/generateText call previously read model: resolvedPrompt?.model ? model : model — a dead ternary (both branches return the same closure-captured LanguageModel) that silently discarded resolvedPrompt.model (the per-prompt model trigger.dev's own prompts.define() declares, e.g. gpt-4o-mini for detect-highlights-content-type) in all 7 places. No step could ever actually run on a model different from the global default. Fixed by provider-factory.ts::createLlm() building an LlmModelBag ({ default, chapters?, highlights?, contentTypeDetection?, smartMetadata?, translation? }, via the extracted buildLanguageModel(config, modelId) — one LanguageModel per configured override, cheap since it's plain object construction and createLlm() runs once per task) instead of a single LanguageModel; createAiSdkLlmProvider(models: LlmModelBag) now picks models.<step> ?? models.default per method. resolvedPrompt.model is deliberately no longer read anywhere — trigger.dev's own prompt-model field lives in a separate admin surface this app doesn't control and would just drift against the real (env+DB) config; resolvedPrompt.text/.config.temperature are unaffected and still wired exactly as before. Unit-tested in packages/ai-pipeline/src/provider-factory.test.ts (env/DB merge + one case per llmProvider branch) and providers/ai-sdk-llm.test.ts (mocks ai's generateObject/generateText, asserts each method picks the right model — the regression guard for the dead-ternary bug specifically).
Frontend Realtime (apps/dashboard, trigger.dev Realtime) — VideoDetailPage subscribes to the asset's live transcode-orchestration/ai-pipeline run status instead of relying solely on its 4s interval poll. jobs.trigger_run_id/ai_jobs.trigger_run_id (set on triggerTask()'s return value in services/process-asset.ts/services/node-reaper.ts, and via ctx.run.id inside ai-pipeline.ts itself, since that task creates its own ai_jobs row) correlate each DB row to its trigger.dev run. GET /v1/assets/:id/realtime-token (routes/assets.ts) mints a scoped, 6h-expiry read-only Public Access Token — trigger-client.ts::mintPublicAccessToken reimplements auth.createPublicToken() for a root secret key locally (one POST /api/v1/auth/jwt/claims call + a local jose HS256 sign, mirroring @trigger.dev/sdk's own auth.js), so apps/api-edge still never depends on @trigger.dev/sdk. The dashboard's AssetRealtimeWatcher component (@trigger.dev/react-hooks's useRealtimeRun, one subscription per run id) calls back into VideoDetailPage on any run-status change, which then slows its own fallback poll from 4s to 20s (POLL_INTERVAL_WITH_REALTIME) — Realtime becomes the primary update path, the interval poll becomes a safety net (kept for assets with no trigger_run_id, e.g. pre-cutover jobs).
Build order
@strum-vod/db must be built first — apps/transcoder and apps/api-edge depend on it. The root pnpm run build handles this via workspace ordering.
Architecture
Browser
├── Dashboard (CF Workers + Assets) → API (Fly.io / Docker)
└── Player app (CF Pages) → API (Fly.io / Docker)
Upload paths:
Standard: Browser → API (presigned URL) → R2
Resumable: Browser → API (TUS, @tus/server + @tus/s3-store) → R2
Transcode (QStash push, two deployables — see "Worker architecture" below):
API → QStash (HTTP push, /qstash/jobs) → Transcoder (Go) → ffmpeg ladder + upload → R2
→ QStash (HTTP push, /qstash/ai) → Worker (Node) → AI pipeline
AI transcription (sync providers — local/deepgram):
/qstash/ai (HTTP) → Worker → whisper/LLM → upload → R2
AI transcription (async provider — modal, scale-to-zero GPU):
/qstash/ai (HTTP) → Worker → POST Modal /transcribe → 202 {callId} [handler returns without blocking]
→ Modal: cold start + inference in a spawned call (no long HTTP hold)
→ POST <API_PUBLIC_URL>/v1/ai/whisper-callback [HMAC-signed]
→ apps/api → QStash /qstash/ai (resume-ai-pipeline | retry-transcription) → Worker resumes
a QStash cron (ai-reaper, every 5 min) re-dispatches or fails modal jobs whose callback
never arrived (transcriptionAttempts gate). See routes/ai.ts
Audio stem separation (Demucs — async modal only, same pattern as async transcription):
/qstash/ai (HTTP) → Worker → POST Modal Demucs /separate → 202 {callId} [fire-and-forget]
→ Modal (infra/modal-demucs/app.py): Demucs 2|4 stems → upload to R2 (playback/{id}/ai/stems/)
→ POST <API_PUBLIC_URL>/v1/ai/demucs-callback [HMAC-signed, DEMUCS_WEBHOOK_SECRET]
→ apps/api → marks separationStatus COMPLETED | publishes retry-separation (MAX_SEPARATION_ATTEMPTS)
the ai-reaper cron also reaps stuck separations (SEPARATION_STALLED_MS gate). See routes/ai.tsHistorical — describes the fully deleted pre-trigger.dev dispatch path. The flowcharts below show the QStash-push transcode + AI pipeline that
apps/api(Fastify) /apps/worker's/qstash/aiconsumer used to run. Both are gone (2026-08-16/17):apps/apiis deleted,apps/workerno longer runs any AI code (thehandleAiJobconsumer, the AI worker and its tests were removed — see "Worker architecture" and "Orchestration migration" below), and the AI pipeline now lives inapps/trigger-taskswith dispatch owned byapps/api-edge. QStash survives only forapps/worker's analytics/backup/org-delete crons. Kept only because the async Modal webhook contract below is still accurate for the trigger.dev tasks that call Modal.
The full transcode → AI → webhook pipeline as a flowchart (QStash replaces BullMQ queues + Redis Streams; Redis remains only for metering, live presence and worker heartbeats):
The async Modal transcription flow as a sequence diagram (same as infra/modal-whisper/README.md):
Worker architecture: Go transcoder + Node AI, joined by QStash
Historical — describes a fully deleted dispatch path (2026-08-16/17).
apps/api(the Fastify app that published to/qstash/jobs,enqueuePlatformTranscode,prewarmFleet) is deleted entirely.apps/transcoder'sinternal/qstashHTTP job-consumer is deleted too —cmd/transcoderis pull/claim only now, no inbound HTTP job intake at all. The legacy AI path described below is deleted as well — theapps/workerAI worker (handleAiJob,/qstash/ai, the ai-reaper cron,ai/pipeline, the AI e2e/unit tests) was removed together with the trigger.dev cutover; the AI pipeline is fully owned byapps/trigger-tasks, and the Modal async contract is consumed by those tasks' providers, not byapps/worker. Current dispatch:apps/api-edge'senqueueTranscodealways triggers thetranscode-orchestrationtrigger.dev task, which wakeswakeStrategy=fly_apimachines directly via the Fly Machines API and waits forPOST /v1/worker-agent/jobs/:id/complete— see "Worker unification" and "Orchestration migration" below for what's actually live. QStash still exists, just nowhere in the transcode path — only forapps/worker's analytics/backup/org-delete crons (triggered byapps/api-edge's Cloudflare Cron Triggers, not by anything described below). The rest of this section (AI dispatch, Modal webhook contract, Fly wake/self-stop mechanics) is kept for the parts still accurate — the Modal async contract hasn't changed, only who runs the pipeline steps.
The transcode pipeline is split across two deployables — porting the ffmpeg-heavy core to Go for lower memory/CPU overhead while keeping the AI pipeline (LLM/Whisper calls, I/O-bound, no Go SDK equivalent for schema-enforced generation) in Node. QStash (HTTP push + cron) replaced the old BullMQ queues + Redis Streams Node↔Go boundary (see docs/dev/qstash-migration-plan.md): every job hop is now a signed HTTP POST between the deployables, with QStash owning retries, dedup (Deduplication-Id) and the cron schedules. Redis remains only for the low-level KV uses: metering counters, live presence, and worker heartbeats. The old apps/worker/src/{bridge.ts,streams.ts} and apps/transcoder/internal/queue are deleted.
apps/apipublishes transcode jobs directly to QStash → the transcoder's/qstash/jobs(services/process-asset.ts::enqueuePlatformTranscode, shared by every call site; dedupjobId,Upstash-Delay~5s for cold start, timeout 2h, retries 5).POST /v1/assets/:id/processalso wakes the transcoder machine viaprewarmFleet(see below).apps/transcoder(Go) consumes jobs over HTTP (internal/qstash/handler.go:Upstash-Signatureverify → semaphore athw.Concurrency→ 429 when full (QStash holds + retries) →ProcessMessage). On success it publishes ai-dispatch toWORKER_PUBLIC_URL/qstash/aiand webhook-dispatch toAPI_PUBLIC_URL/qstash/webhook(dispatchAiProcessing/dispatchDomainEvent, both with dedupe ids and QStash-owned retries).apps/worker(Node,strum-vod-worker) — AI + analytics HTTP consumers. Never touches ffmpeg for the main ladder anymore (still shells out toffmpegfor AI highlight-clip cutting,ai/highlights.ts). The routes insrc/qstash-routes.tsverify the signature and dispatch to shared cores:- AI worker (
src/ai-worker.ts::handleAiJob) — runs onPOST /qstash/ai(job shape carried by thejobTypefield). Downloads the pre-extractedai/audio.mp3and the archived original source (sources/<assetId>/input.mp4, both uploaded by the Go transcoder) from R2, then runsai/process.ts(Whisper/Deepgram transcription, LLM chapters/highlights). Silent sources are skipped entirely — the transcoder leavesassets.ai_audio_pathNULL when the source has no audio track, and the worker marks theai_jobsrow + all four steps SKIPPED instead of failing the audio download. ThejobTypefield carries the job shapes:process(main flow) — creates theai_jobsrow, runsprocessAi(). For sync providers (local/deepgram/assemblyai) this blocks until transcription is done, then runs subtitles/chapters/highlights. For the asyncmodalprovider it only dispatches (see below):processAiPOSTs the audio to Modal, stores thecallIdinproviderJobId, and returns — the handler "returns" until the webhook callback lands.resume-ai-pipeline— published byapps/api's/v1/ai/whisper-callbackwhen a modal transcription completes. Reuses the existingai_jobsrow (transcript already written by the callback) and runs only the post-transcription steps (resumeAiPipeline(): subtitles, chapters, highlights).retry-transcription/retry-transcription-sync— published by the callback when Modal reports a failed transcription; re-dispatches the audio (or falls to the next sync chain provider).reap-stuck-transcriptions/reap-stuck-separations— a QStash cron (strum-vod-ai-reaper, every 5 min) hits/qstash/cron/ai-reaper;reapStuckTranscriptions()re-dispatches (or fails) modal jobs whose callback never arrived, gated by an atomictranscriptionAttemptsincrement so concurrent reapers can't double-dispatch.
- Analytics aggregation (
src/analytics-worker.ts::runAnalyticsJob) — the three QStash crons (strum-vod-analytics-{hourly,daily,cleanup}) hit/qstash/cron/analyticswith{type}in the body. - DB backups (
src/db-backup.ts::runDbBackup) —POST /qstash/backup(daily QStash cron + manual trigger). Org deletes (src/org-delete.ts::processOrgDelete) —POST /qstash/org-delete, published by the API's admin route.
Async (modal) transcription contract —
infra/modal-whisper/app.pyis a scale-to-zero Modal app that never holds a request open across cold start + inference: the worker POSTsmultipart/form-data(file,callback_url,asset_id,ai_job_id, plus an optional ISO-639-1languagehint forwarded to faster-whisper'stranscribe(language=...)to bias detection — the forced language is still applied at callback time) to<endpoint>/transcribewithAuthorization: Bearer <WHISPER_API_KEY>and gets an immediate202 {callId}; the actual transcription runs in a.spawn()ed call, which later POSTs the result to<API_PUBLIC_URL>/v1/ai/whisper-callbackwith anX-Signature: sha256=<hmac>header over the raw body (secret:WHISPER_WEBHOOK_SECRET).apps/apiverifies the HMAC withtimingSafeEqual, correlatescallId === provider_job_id(anti-replay), and is idempotent for already-settled jobs. Seeinfra/modal-whisper/README.mdfor the full deploy/setup.Async Demucs separation contract —
infra/modal-demucs/app.pymirrors the whisper app: the dispatcher (packages/ai-pipeline/src/providers/demucs.ts'sdispatchModalDemucs(), called from theseparate-audio-stemstrigger.dev task) POSTsmultipart/form-data(file,callback_url,asset_id,ai_job_id,stems2|4) to<endpoint>/separatewithAuthorization: Bearer <endpoint's apiKey>and gets202 {callId}; the spawned call runs Demucs (htdemucs_ft2 stems → vocals/no_vocals,htdemucs4 stems → vocals/drums/bass/other), uploads each stem MP3 to R2 atplayback/{assetId}/ai/stems/{name}.mp3, then POSTs the signed result toPOST /v1/ai/demucs-callback(DEMUCS_WEBHOOK_SECRET). Non-fatal to the pipeline: dispatch failure only marksseparationStatusFAILED and transcription continues. No single-URL config exists — a Demucs endpoint is configured exclusively through the dashboard's endpoint pool (/admin/ai→ "Pool de Endpoints",AiProviderSettings.separationEndpoints), tried in priority order with failover to the next entry on error/timeout (dispatchModalDemucs()). Seeinfra/modal-demucs/README.mdfor the full deploy/setup.- AI worker (
apps/transcoder(Go,strum-vod-transcoder) — consumes jobs via the HTTP handlerinternal/qstash(see above), resolves the source (local shared volume → yt-dlp URL import → S3 fallback, same 3-tier priority as before), runs the full HLS ladder as a single ffmpeg pass (internal/ladder, 7 renditions 360p–4320p + one shared EXT-X-MEDIA audio track) + thumbnails + audio extraction viaos/exec+ ffmpeg/ffprobe/yt-dlp, uploads everything to R2, marks the assetready, then publishes AI + webhook dispatches via QStash and archives the source. Scales to zero (fly.transcoder.toml,min_machines_running=0) via a split wake/stop design, since Fly's own proxy-idleauto_stop_machinesis unsafe for a job-driven process (it could stop a machine mid-transcode, since it has no HTTP connection open while encoding):- Wake —
auto_start_machines=true; the API wakes the transcoder in parallel with the worker on every job enqueue (apps/api/src/routes/assets.ts::prewarmFleet→services/fly-machine.ts::startTranscoderMachine), using the Fly Machines API (POST /machines/:id/start, needsFLY_API_TOKEN) or falling back to the publichttps://<FLY_TRANSCODER_APP>.fly.dev/wakeproxy. The QStash delivery itself also autostarts the machine (Fly's proxy autostarts on the inbound HTTP request); the transcoder's/wakehandler (internal/health/health.go) just acks. - Stop — self-managed by
internal/selfstop(Go), not by Fly's proxy: it tracks in-process job activity viaMarkBusy/MarkIdle(wrapped around every job incmd/transcoder/main.go) and, once it has observed zero active jobs continuously forSELF_STOP_IDLE_SECONDS(default 120s), just triggers an ordinary graceful shutdown (the samecontext.CancelFuncSIGTERM already drives) so the process exits 0. No Fly Machines API call or token required — a Fly Machine automatically transitions tostoppedwhen its init process exits on its own (see Fly's long-running-tasks blueprint). Gated onFLY_MACHINE_ID(auto-injected by Fly) so it never fires in local dev/Docker Compose.
- Wake —
The interop boundary is QStash (HTTP push): the token/signing keys are required on the transcoder process (env/config validation fails at boot without them), with TRANSCODER_PUBLIC_URL and API_PUBLIC_URL as the publish targets (see .env.example). The Node apps/worker has been deleted — its cron jobs (analytics/backup/org-delete) are now trigger.dev scheduled tasks. Enum/constant strings (ASSET_STATUS, JOB_STATUS, S3_PATHS, etc.) are mirrored by hand in apps/transcoder/internal/constants from packages/db/src/constants.ts — there is no generated single source yet, so a change to one needs the other updated too. apps/transcoder/internal/constants/parity_test.go enforces the mirror (parses constants.ts and fails on drift); run it with pnpm test:transcoder (uses -count=1 since Go's test cache can't see changes to the .ts file).
Worker unification (done) — one fleet, pull/claim, no QStash in the transcode path
The "Worker architecture" section above is historical — the QStash-push dispatch path it describes (enqueuePlatformTranscode, internal/qstash, internal/pipeline) has been fully replaced, its code deleted from every app that had it. Full original design (background/rationale, not current status): /home/juninho/.claude/plans/planeje-structured-fog.md.
What changed and why: three separate "compute worker" systems (the Fly platform fleet pushed via QStash + direct DB/R2 credentials; customer-facing self-hosted "Nodes"; admin-managed "Platform Workers") are being collapsed into one — there is no third-party BYO-compute customer, so "Nodes" was pure duplication of "Platform Workers." Target: one Go binary (apps/transcoder/cmd/transcoder, content-identical to cmd/node-agent on purpose during the transition), one broker (apps/api-edge/src/routes/internal/worker-agent.ts, /v1/worker-agent/*), one credential type (mt_live_... machine tokens, machine_tokens/machines/dispatch_jobs tables), pull/claim for every worker including the Fly fleet — eliminating QStash from the transcode dispatch path entirely. The remaining analytics/backup/org-delete crons have also been migrated to trigger.dev scheduled tasks, so QStash is no longer used for any cron jobs.
Machine-token leases (machine_tokens.expires_at): tokens are long-lived by default (manual revoke only), but POST /v1/admin/machine-tokens accepts an optional ttlDays (1–365) that sets an expires_at lease. requireWorkerAuth (apps/api-edge/src/middleware/auth.ts) rejects any token past its expiry at request time — even without a manual revoke — so a leaked/stolen throwaway token's blast radius is bounded. Default to a TTL for ephemeral workers (laptops, burst boxes); keep long-lived tokens only for the always-on fleet. Dashboard /admin/machines exposes the TTL when creating and shows "expirado"/"expira em" status per token.
Status — LIVE in production, verified end to end with real infra (2026-08-16):
enqueueTranscode(apps/api-edge/src/services/process-asset.ts) unconditionally triggers thetranscode-orchestrationtrigger.dev task (apps/trigger-tasks) — no more fly_only/nodeRouting branch, no rollout flag.transcode-orchestrationinserts adispatch_jobsrow, wakes anymachinesrows withwakeStrategy='fly_api'for real (apps/trigger-tasks/src/lib/fly-wake.ts, direct Fly Machines API call — not a placeholder log line), creates a wait token, and waits.- Any worker running the unified binary polls
GET /v1/worker-agent/jobs/claim(no org/scope filter — firstqueuedrow wins,FOR UPDATE SKIP LOCKED);/completetriggers theai-pipelinetask directly, emits the customer webhook in-process, and resolves the wait token — no QStash hop anywhere in this path. apps/transcoder/internal/nodeagent/agent.gohandles all 4 job types:transcode,extract_audio,mix_stems,render_highlights(the last one cuts clips fordiscover-highlights's LLM-discovered suggestions — see the "Orchestration migration" section above). Stem downloads formix_stemsuse presigned GET URLs minted by the claim handler (worker-agent.ts'sstemUrls), not a generic download-by-key endpoint.- Production deploy done:
strum-vod-api-edgeredeployed to Cloudflare Workers with/v1/worker-agent/*+/v1/admin/machine-tokenslive,TRIGGER_SECRET_KEYset, trigger.dev prod fully provisioned and deployed (see "Orchestration migration" above). The production Postgres (Neon) was missing a large amount of schema — not just this migration's tables, but also Better Auth columns/tables (users.email_verified/image,sessions/accounts/verifications/invitations) and the older platform-worker fleet tables — apparently never migrated since those features were built. Fixed by running the real migration for real:DATABASE_URL=<prod> ... pnpm --filter @strum-vod/api run db:migrate(supply dummy values for the other required env vars —db-migrate.tsonly touches the DB). Schema changes now migrate automatically in CI:.github/workflows/deploy-backend.ymlrunspnpm db:migrate(the env'sDATABASE_URLGitHub secret) before every backend deploy, so landing apackages/dbchange reaches the database without a manual step. Verified with a real asset: upload →process→ a worker binary running on an operator's own machine (not Fly) claimed it → real ffmpeg encode →ready→ AI pipeline ran for real (Deepgram transcription + subtitles). - OLD systems and
internal/pipeline/internal/qstash(Go) — fully deleted (2026-08-16):routes/node-agent.ts,routes/nodes.ts,routes/platform-workers.ts,services/nodes.ts,cmd/node-agent, thenk_/pw_token families, and thenodes/node_jobs/platform_workers/platform_worker_tokenstables — deleted fromapps/api-edge(and, along with the rest of it, fromapps/api, which no longer exists at all — see below).POST /v1/assets/:id/extract-audio//stems/mix(routes/assets.ts) now dispatch throughtranscode-orchestration/dispatch_jobssame as full transcode — the Go side (internal/nodeagent/agent.go) already handled all 4 job types before this cutover, so no worker-side gap.apps/transcoder/go.moddropped from 6 direct deps to 1 (golang.org/x/sync) once the credentialed DB/S3/QStash packages had no callers left. - Not yet done: deploying the unified binary to the Fly fleet itself (
wakeStrategy=fly_api) — today it only runs pull/claim on an operator's own machine, which was always the point but means Fly'sstrum-vod-transcoderapp is currently idle/unused. apps/api(Fly, Fastify) — fully deleted (2026-08-16), not just superseded. Its Fly app (strum-vod-api) had already been decommissioned independently of this cleanup — confirmed via a realfly deploy("app not found") and a DNS lookup onstrum-vod-api.fly.dev(doesn't resolve) — so nothing live depended on it. Deleted: the wholeapps/api/directory (routes, services, middleware, Dockerfiles,fly.toml), the strangler-fallback proxy inapps/api-edge(proxy.ts,FLY_API_ORIGIN,app.all('*', proxyToFly)— replaced with a plain 404 handler),apps/cli's E2E suite (it spawned a realapps/apisubprocess via Testcontainers — no replacement yet, a real coverage gap),packages/api-client(an empty, zero-consumer orval scaffold that only ever existed to consumeapps/api's never-actually-generated OpenAPI schema), and the non-edgedocker-compose.ymlservices (api/worker/transcoder/dashboardcontainers built around the old direct-DB-credentialed architecture —pnpm devnow meanspnpm dev:edge:full, the only working local dev stack). The schema migration runner moved:apps/api/src/db.ts'srunMigrations()/bootstrapDefaultOrg()is nowpackages/db/src/migrate.ts(pnpm db:migratefrom the repo root, orpnpm --filter @strum-vod/db run migrate) — same idempotent raw-SQL migrations, same DATABASE_URL-only contract, verified against real production Postgres before the cutover. Known real gaps left by this deletion, not yet addressed: (1)apps/api-edgehas zero automated test coverage of its own beyondtypecheck— apps/api's fine-grained Testcontainers E2E suite (auth, assets, billing, orgs, permissions, settings, webhooks, ...) is gone and was never a suite that actually exercised api-edge's ported code; (2) the monthlycost_snapshotscron (apps/api's own/qstash/cron/cost-snapshot+ itsregisterQstashSchedulesentry) has no home anymore — reads still work (apps/api-edge/src/services/costs.ts), new snapshots don't get generated; (3) the legacy AI e2e tests (ai-worker.test.ts,full-pipeline.test.ts) were deleted with the legacy AI path (2026-08-17) — the Go-side transcode is now covered byapps/transcoder's own integration tests and the AI pipeline by trigger.dev, but there's no automated test asserting the api-edge → trigger.dev → worker-agent → Go end-to-end chain.
Cloudflare infrastructure (wrangler)
Cloudflare resources are deployed with wrangler — no IaC tool is used anymore.
Cloudflare resources:
- R2 bucket
strum-vod(video/audio storage) +strum-vod-backups(DB backups) — CORS is applied frominfra/r2/strum-vod-cors.jsonviapnpm r2:cors:set(required for browser-direct presigned uploads from the dashboard/player) - Dashboard as a Cloudflare Worker + static assets (SPA mode) —
apps/dashboard/wrangler.toml - Player as a Cloudflare Pages app —
apps/player/wrangler.toml(pages_build_output_dir = "dist")
Resumable video upload is served by apps/tus-edge, a separate Cloudflare Worker (TUS protocol) — not part of the apps/api-edge Worker itself.
Monorepo using pnpm workspaces, 17 packages/apps (9 apps + 8 shared packages):
apps/api-edge— Hono REST server on Cloudflare Workers. The one and only backend —apps/api(Fastify, Fly) is deleted entirely (2026-08-16), see "Worker unification" above for the deletion writeup. Env comes fromwrangler.toml[vars]+wrangler secret, no Zod schema (Workers'c.envis already typed per-binding). DB access isdrizzle-orm/neon-http(src/db.ts) — stateless per request, no connection pool to manage, cannot speak plain TCP Postgres (seedocs/dev/LOCAL_EDGE_DEV.md'sneon-proxycontainer for how local dev bridges this). Routes live undersrc/routes/{customer,console,internal}/*.ts, split by who calls them, not by resource:customer/— the video-resource API (assets, collections, playback, comments, webhooks, ai). Acceptsmk_API keys and a dashboard session — an org's own integration and the dashboard's asset-management UI hit the exact same routes.console/— dashboard-only account/platform management (auth, orgs, billing, settings, backups, stats, diagnostics, admin, machines).requireSession()(middleware/auth.ts) rejectsmk_keys on all of these except diagnostics/admin/machines, which are alreadyuserId-gated viarequireSuperadmin().internal/— never reached through the dashboard's typed RPC client: the unified worker fleet (worker-agent.ts, machine tokens), infra liveness probes (health.ts), the operator install script, and signed provider callbacks (Modal'sai-callbacks.ts, Stripe'sbilling-webhook.ts).
index.tscomposescustomer/+console/into onedashboardRoutessub-app and exportsexport type AppType = typeof dashboardRoutes—packages/api-client'shc<AppType>()(a typed Hono RPC client, not Orval/OpenAPI) types the dashboard's API calls off this directly, sointernal/routes never leak into that client's surface.internal/routes are mounted afterward onto the same running app for actual serving (still one Worker, onefetch), just outsideAppType. A 404 JSON handler catches anything unmatched (no more strangler-fallback proxy, nothing to fall through to).openapi/generate-customer-spec.ts(pnpm --filter @strum-vod/api-edge run openapi:customer, oropenapi:customer:allto also regeneratecustomer.d.tsviaopenapi-typescript) generatesopenapi/customer.json(OAS 3.1) for external customers — explicitregisterPath()calls against Zod schemas imported directly from@strum-vod/api-contracts(assets,playback,organization,commonnamespaces), not a source-scan, so schema drift shows up as a type error in the generator script itself rather than silently. Most paths carry a realresponseSchemaoff the imported contracts; a handful of simple/legacy endpoints still fall back to the generic{ data: unknown }shape (GENERIC_RESPONSE) — see the script's own file comment.console/andinternal/are not covered — this is customer-surface only, and the generator file carries an explicit scope-note comment abovemain()to that effect: a prior version of this file (pre-2026-08-24) leakedroutes/console/*.ts's auth/orgs/settings/billing/ admin/stats/diagnostics paths into this spec, plus three/v1/ai/*paths that didn't correspond to any real route at all (the real AI-config surface is/v1/admin/settings/ai/*, console-only) — fixed by scopingregisterPath()calls back to whatroutes/customer/*.tsactually mounts (down from 100+ registered operations to the real 43 paths / 52 operations). If you add a newroutes/customer/*.tsendpoint, add itsregisterPath()call here too — nothing generates this file's route list automatically.test/unit/routes-access-guard.test.tssource-scansroutes/console/*.tsthe same way, failing if a route callsrequireAuth()without also chainingrequireSession()/requireSuperadmin()(an explicit allowlist coversGET /v1/config, the one deliberate exception).MCP server, transcode dispatch (
enqueueTranscode), the unified worker broker (/v1/worker-agent/*), and the trigger.dev Realtime token minting are all documented in their own sections elsewhere in this file — this bullet is deliberately not a full file-by-file catalog (api-edge's structure is covered piecemeal throughout "Worker unification"/"Orchestration migration" above and the Critical Rules below) to avoid drifting out of sync with two descriptions of the same code.- Known gap: no automated test suite of its own beyond
typecheck— see the "apps/api deleted" bullet above for what coverage was lost and hasn't been replaced yet.
apps/trigger-tasks— trigger.dev tasks: AI pipeline (transcription, chapters, highlights, separation, translation), transcode orchestration, cron maintenance (db-backup, reclaim-stale-ai-jobs, reclaim-stale-dispatch-jobs), org-delete, webhook delivery, and 4 background agent tasks (src/trigger/agents/). The legacyapps/workerwas deleted (2026-08-17); its crons are now trigger.dev scheduled tasks. The shared provider unit tests moved topackages/ai-pipeline. (Theanalytics-hourly/analytics-daily/analytics-cleanupfiles mentioned by earlier docs no longer exist — that surface was superseded byreclaim-stale-ai-jobs/reclaim-stale-dispatch-jobs/db-backup.)src/trigger/— one file per task:ai-pipeline.ts,ai-pipeline-continue.ts,transcode-orchestration.ts,separate-audio-stems.ts,transcribe-audio.ts,generate-subtitles.ts,generate-chapters.ts,generate-smart-metadata.ts,discover-highlights.ts,score-highlights.ts,translate-subtitles.ts,analyze-transcript.ts,db-backup.ts,db-backup-manual.ts,org-delete.ts,process-webhook-delivery.ts,reclaim-stale-ai-jobs.ts,reclaim-stale-dispatch-jobs.ts,agents/{video-qna,metadata-refinement,translation-assistant,transcript-insights}-agent.tsschemaTask+metadataeverywhere (2026-08-24): every task in this app isschemaTask()(zod-validated payload,TaskPayloadParsedErroron a bad shape instead of running with garbage) —task()survives only forschedules.task()-based crons (db-backup,reclaim-stale-ai-jobs,reclaim-stale-dispatch-jobs), whose payload shape is trigger.dev's own fixedScheduledTaskPayload, not user-definable. Shared payload shapes ({assetId, aiJobId},AiOptions) live insrc/lib/schemas.ts(AssetAiJobPayloadSchema,AiOptionsSchema) rather than being redefined per file. Every task also callsmetadata.set(...)at its real lifecycle transitions (mirroring what was alreadyconsole.logged, not inventing new status vocabulary) —status/phase/errorare the common keys, feedingruns.subscribeToRun/the dashboard's existingAssetRealtimeWatcherRealtime consumer for free. For the 4 tasks driven bysrc/lib/ai-task.ts's sharedrunAiStep/loadAiStepContext(generate-subtitles,generate-chapters,generate-smart-metadata, anddiscover-highlights's skip path), the metadata calls live centrally in that helper, not duplicated per call site. Test gotcha:metadata's real implementation throwsMethod not implementedoutside an actual trigger.dev run context (aNoopRunMetadataManager, not a silent no-op as the SDK docs imply) — every test file that imports a task module transitively touchingmetadata.setmust mock@trigger.dev/sdk'smetadataexport as a chainable no-op stub (see anyvi.mock('@trigger.dev/sdk', ...)block insrc/trigger/__tests__/*.test.tsfor the pattern).src/lib/— shared helpers:db.ts(drizzle client),s3.ts(R2 upload/download),ai-config.ts(AI provider config resolution),ai-task.ts(sharedrunAiStep/loadAiStepContextfor the four AI steps above),schemas.ts(sharedschemaTaskpayload schemas),db-backup.ts(pg_dump → R2)trigger.config.ts— project config: PGDG build layer for pg_dump 18,syncEnvVars+ static vars (S3_BUCKET,S3_BACKUP_BUCKET,TRANSCRIPTION_PROVIDER), DEPLOY_SYNCED_VARS
apps/transcoder— Go: the ffmpeg-heavy transcode ladder (ported from the oldapps/workerfor lower memory/CPU overhead — see "Worker architecture" above). Hardware-adaptive: auto-detects CPU/RAM (cgroup-aware) at startup.cmd/transcoder/main.go— Entry point: config, DB pool, Redis (metering/heartbeat), QStash handler wiring, health server, graceful shutdowninternal/config— Env parsing + hardware-adaptive concurrency/thread/pool-size formulasinternal/qstash— HTTP job consumer (/qstash/jobs):Upstash-Signatureverify, semaphore athw.Concurrency+ 429,ProcessMessageunder the app context; also the publish client (QstashPublisher) used by the pipeline for ai/webhook dispatch.internal/queue(the Redis Streams consumer) was deleted (Phase 4)internal/pipeline— Job orchestration: source resolution, single-pass ladder + per-profile MP4 remux, audio/thumbnails, upload, QStash AI/webhook dispatchinternal/ladder— Rendition profiles (Ladder), single-pass multi-output ffmpeg command, shared EXT-X-MEDIA audio, master playlist, MP4 remuxinternal/vaapi,internal/thumbnails,internal/ytdlp,internal/ffmpegx— Hardware-accel args, tiled sprite/VTT generation, yt-dlp wrapper, ffmpeg/ffprobe process spawninginternal/storage— S3/R2 client: streaming upload/download, directory-tree bulk uploaderinternal/db— pgx pool + hand-written queries againstpackages/db/src/schema.ts's tablesinternal/constants— Hand-mirrored copy ofpackages/db/src/constants.ts's enum strings (keep both in sync manually)
apps/dashboard— React SPA (Vite + Tailwind CSS v4). Authenticated management UI for uploading and managing assets. PWA — installable (manifest + Workbox SW viavite-plugin-pwa, see "PWA" below).src/App.tsx— Entry point with routing and layout (no embed/watch routes — those are in the player app)src/pages/NewVideoPage.tsx— Upload page with TUS resumable / presigned-URL togglesrc/components/Player.tsx— Full-featured HLS player (used inside the dashboard for preview)src/components/InstallAppButton.tsx— "Add to Home Screen" button (nativebeforeinstallprompton Android/Chrome, instructions modal on iOS)src/lib/pwa.ts—usePwaInstall()hook (beforeinstallprompt / standalone / iOS detection)src/lib/types.ts— TypeScript interfacessrc/lib/api.ts— Authenticated API client helpersrc/lib/helpers.ts— Utility functions (timeAgo, formatDuration, STATUS_CFG, etc.)
apps/player— Standalone Cloudflare Pages app (Vite + React). Serves the public embeddable player. No auth, no dashboard chrome. Deployed with wrangler pages (pnpm deploy:player→wrangler pages deploy dist --project-name=player). PWA — installable too (samevite-plugin-pwasetup as the dashboard).src/App.tsx— Routes:/embed/:playbackId,/watch/:playbackIdsrc/components/Player.tsx— HLS player (ported from dashboard, no i18n dependency)src/components/EmbedPlayer.tsx— Embed wrapper (fetches/v1/playback/:id)src/components/WatchPage.tsx— Minimal public watch pagesrc/lib/analytics.ts— PlayerAnalytics class (heartbeat + batch flush)src/lib/api.ts— Unauthenticated API client
PWA (dashboard + player)
Both frontend apps are installable PWAs (vite-plugin-pwa v1 + Workbox generateSW):
- Config —
VitePWA()in eachvite.config.ts:injectRegister: false(SW registered from React, see below),display: standalone,navigateFallback: '/index.html'for the SPA routes,skipWaiting/clientsClaim, 5 MiBmaximumFileSizeToCacheInBytes(the dashboard bundle is big). Update mode differs per app: the player usesregisterType: 'autoUpdate'(silent takeover); the dashboard usesregisterType: 'prompt'and surfaces updates through a "New version available → Reload" toast (src/components/UpdateToast.tsx,useRegisterSWfromvirtual:pwa-register/react) so an update lands like a native app's update prompt. - Registration is top-frame-only —
UpdateToastrenders nothing whenwindow.self !== window.top, so the SW registers only in top-level frames. Chrome disallows SW registration in cross-origin iframes, and the dashboard serves/embed//watchpages that are embedded on other sites; we don't want an embed to claim the origin's SW. - Icons — generated by
scripts/generate-pwa-icons.py(Pillow; run it after changing the brand color):public/pwa-192x192.png,pwa-512x512.png,pwa-512x512-maskable.png,apple-touch-icon.png,favicon.svgin each app. Manifest icons are referenced frompublic/. public/_headers— both apps ship a_headersfile so Cloudflare (Workers Assets for the dashboard, Pages for the player) servessw.js/workbox-*.js/manifest.webmanifestwithCache-Control: no-cache— a long edge cache would pin clients to a stale service worker and break PWA updates. Do not remove it.- Native-app feel —
viewport-fit=cover+pt-safe/pb-safe/px-safeutilities (env(safe-area-inset-*)) keep content clear of the notch/home indicator when installed;overscroll-behavior-y: none,-webkit-tap-highlight-color: transparentandtouch-action: manipulationkill web-page-isms (rubber-banding, tap flash, double-tap zoom); layouts useh-dvh/min-h-dvhinstead of100vhso the mobile URL bar doesn't cause layout jumps.
Documentation site (VitePress)
The docs/ directory doubles as a VitePress site (docs/.vitepress/config.mts) — the same markdown that lives in the repo is served as a browsable docs site. Wrapper pages (docs/readme.md, docs/docker.md, docs/contributing.md, docs/security.md, docs/changelog.md, docs/claude.md) import the root-level markdown (../README.md, ../DOCKER.md, etc.), so there is one source of truth — edit the root files, the site updates.
Config —
docs/.vitepress/config.mts: dark theme,vitepress-plugin-mermaid(mermaid blocks from CLAUDE.md/architecture render inline), local search,ignoreDeadLinksfor repo-relative links, and abuildEndhook that copiesdocs/diagrams/(SVG/PNG) into the build output.Commands —
pnpm docs:dev(dev server, :5173),pnpm docs:build(static build →docs/.vitepress/dist),pnpm docs:preview(serve built site, :4173),pnpm deploy:docs(build +wrangler pages deploy).Deploy — Cloudflare Pages project
strum-vod-docs(seedocs/wrangler.toml)..github/workflows/docs.ymlrebuilds + redeploys automatically on every push tomaintouchingdocs/**or the root markdown the site imports; PRs get a preview URL. Secrets:CLOUDFLARE_API_TOKEN,CLOUDFLARE_ACCOUNT_ID.VITE_DOCS_URL— dashboard build-time env var for the sidebar's "Documentation" link (apps/dashboard/src/components/layout/Sidebar.tsx), which opens in a new tab. Defaults tohttp://localhost:4173(the local preview); set it to the deployed docs URL (e.g.https://docs.strum-vod.dev) in production builds.packages/db— Shared Drizzle ORM schemas (PostgreSQL, pg-core), connection factory, constants.src/schema.ts—collections,assets,renditions,jobs,ai_jobs,highlights,analytics_events,analytics_daily,analytics_asset_stats,settings,comments,reactions,users,organizations,org_members,api_keys,admin_audit_log,domain_events,webhook_deliveries,db_backups,otp_codestables with FK constraints (CASCADE DELETE / SET NULL)src/client.ts—createDb(url, poolConfig?)— PostgreSQL pool factory (postgres-js; SSL forneon.tech/sslmode=require)src/constants.ts— Shared enums (ASSET_STATUS, JOB_STATUS, SOURCE_TYPE, S3_PATHS, ID_LENGTH, COLLECTION_COLORS, TIER_LIMITS, …).isSyncProvider/isAsyncProviderare the canonical sync/async base thatpackages/ai-providers' registrymodemirrors — production branches ontranscriptionMode(); don't delete these predicates (the registry test anchors lockstep)
packages/ai-providers— Shared AI provider registry (single source of truth for transcription + LLM provider metadata):TRANSCRIPTION_PROVIDER_REGISTRY/LLM_PROVIDER_REGISTRYcarry each provider'smode(sync/async),requiredconfig fields,defaultBaseUrl, ahealthProbe(endpoint + auth headers), andfields— JSON-serializable form-field descriptors (ProviderField[]: password/url/select/boolean/group/chain types, i18n label keys, options,visibleForgating) that the dashboard's AI provider form renders from.testProviderHealth()drives the dashboard's "Test connection" flow,resolveTranscriptionChain()is the shared chain resolver, andserializeProviderRegistry()produces the metadata served byGET /v1/admin/settings/ai/providers(healthProbe functions stripped). Consumed byapps/api-edge'sservices/ai-connection-test.ts+ settings route andapps/trigger-tasks/src/lib/ai-config.ts— adding a provider (or editing its form fields) touches only this registry, never the apps. Depends on@strum-vod/dbfor the config types. Data-only: no SDK imports, no HTTP beyond the probe helper.packages/costs— Monthly cost-estimator engine (prices + calculation + DB usage collection, seedocs/costs.md).packages/email— Transactional email templates (OTP magic codes) via Resend / SMTP / consolepackages/subtitles— WebVTT / SRT generation from transcript segmentspackages/player-ui— Shared player components (VideoHeatmap, …)packages/design-tokens— Design tokens / CSS custom properties (brand palette)
Asset Lifecycle
created → (upload-url or upload-token or import) → uploaded → (process) → queued → processing → ready | error
Delete is a hard delete (DELETE /v1/assets/:id removes the DB row — FK CASCADE takes renditions/jobs/ai_jobs — plus the sources/ and playback/ objects in R2). The deleted status constant is retained only to keep the Go-side mirror (apps/transcoder/internal/constants) in sync; no current code path writes it.
Upload Flows
Standard (presigned URL)
POST /v1/assets/:id/upload-url→ API returns a signed S3 PUT URL- Browser PUT the file directly to S3/R2
POST /v1/assets/:id/upload-complete→ API confirms the object exists
Resumable (TUS)
POST /v1/assets/:id/upload-token→apps/api-edgereturns a 15-min JWT (sub: sourceKey, aud: videos, maxLen: fileSize), signed withSHARED_AUTH_SECRET- Browser opens TUS upload to
VITE_TUS_SERVER_URL/upload/videoswithAuthorization: Bearer <token>—VITE_TUS_SERVER_URLpoints atapps/tus-edge, a separate Cloudflare Worker (notapps/api-edgeitself) apps/tus-edge(src/index.ts) verifies the JWT (jose'sjwtVerify, sameSHARED_AUTH_SECRET) on every request — not just creation, so an in-progress upload can't be resumed without a token valid for that exactsourceKey— and streams chunks to R2; the upload id binds tosourceKeydirectly, so the finished object lands at the same keygetSourceKey()expects with no separate move stepPOST /v1/assets/:id/upload-complete→apps/api-edgeconfirms via HeadObject
Key Conventions
- All API responses wrap data in
{ data: {...} }or return{ error: "..." } - Status strings use shared constants from
@strum-vod/db(ASSET_STATUS, JOB_STATUS, etc.) - IDs generated with
nanoid(12), playback IDs withnanoid(16)— lengths defined inID_LENGTH - DB column names use
snake_case, Drizzle schema fields usecamelCase - S3 paths: sources at
sources/{assetId}/input.mp4, HLS output atplayback/{assetId}/— prefixes defined inS3_PATHS - ESM throughout (
"type": "module"in all packages), imports use.jsextensions - No
ACL: 'public-read'on anyPutObjectCommand— R2 uses bucket-level public access, not per-object ACLs - No
PutBucketCorsCommand— R2 CORS is configured via Cloudflare dashboard, not the S3 API - Error handling: API uses centralized error handler (AppError/NotFoundError), Worker wraps DB updates in try/catch
Database
PostgreSQL 16+ (Neon-compatible). apps/api-edge reads/writes via drizzle-orm/neon-http (src/db.ts, stateless per request); the schema migration script (packages/db/src/migrate.ts, pnpm db:migrate) uses drizzle-orm/postgres-js (TCP, hardware-adaptive pool) since it's a plain Node CLI, not a Worker request. Schema defined in packages/db/src/schema.ts. Migrations are raw SQL CREATE TABLE IF NOT EXISTS in packages/db/src/migrate.ts.
Storage
Cloudflare R2 everywhere — both local dev and production. No per-object ACL; CORS and public access are configured at the bucket level with wrangler (see pnpm r2:cors:* and infra/r2/strum-vod-cors.json).
- API (including TUS uploads), transcoder, and dashboard preview all use R2's S3-compatible API (
S3_ENDPOINT=https://<account-id>.r2.cloudflarestorage.com) — same client config everywhere, local dev and production alike. - Buckets are managed with wrangler (
pnpm r2:list). R2 CORS is a bucket-level setting and does not survive a bucket recreate — re-applypnpm r2:cors:setafter creating/replacing thestrum-vodbucket, or browser-direct presigned uploads break (noAccess-Control-Allow-Origin). - Uploads set per-object Cache-Control (
cacheControlForKey()— mirrored inapps/transcoder/internal/storage/s3.go,apps/trigger-tasks/src/lib/s3.ts, and the backfill script). For objects written before that, backfill the headers withpnpm backfill:cache-control(S3 copy-onto-self +MetadataDirective: REPLACE;DRY_RUN=1to preview,PREFIX=...to scope — seescripts/backfill-r2-cache-control.ts).
The S3_PUBLIC_BASE_URL env var controls the URL prefix prepended to HLS manifest paths for public playback.
Docker Build Context Isolation
Each app Dockerfile has a companion Dockerfile.dockerignore that limits the Docker build context to only what that app needs:
| App | Context includes |
|---|---|
apps/dashboard | package.json, pnpm-workspace.yaml, apps/dashboard/ |
apps/transcoder | apps/transcoder/ only — a standalone Go module, no pnpm workspace files needed |
All Dockerfiles use pnpm (via corepack) instead of npm. node_modules and dist directories from the host are excluded to avoid copying broken pnpm symlinks into the build context.
Transcoding Profiles
Defined in apps/transcoder/internal/ladder/ladder.go as ladder.Ladder (hand-mirrored from the retired apps/worker/src/transcoding.ts, which no longer exists — keep the ladder in sync manually; ladder.FilterLadder derives the per-source subset and assigns codecs per tier — see the config table below):
| Quality | Resolution | Bitrate | H.264 profile/level | CodecTag |
|---|---|---|---|---|
| 360p | 640x360 | 1000k | main 3.0 | avc1.4d401e |
| 480p | 854x480 | 1800k | main 3.0 | avc1.4d401e |
| 720p | 1280x720 | 3000k | main 3.1 | avc1.4d401f |
| 1080p | 1920x1080 | 6000k | high 4.0 | avc1.640028 |
| 1440p | 2560x1440 | 10000k | high 5.0 | avc1.640032 |
| 2160p | 3840x2160 | 20000k | high 5.1 | avc1.640033 |
| 4320p | 7680x4320 | 40000k | high 6.0 | avc1.64003c |
FilterLadder keeps only renditions at/below the source's short side (portrait videos skip the taller steps, e.g. a 1080x1920 source only gets up to 1080p) and the configured MAX_RENDITION_HEIGHT cap, always retaining at least the lowest rendition.
Ladder configuration (env):
| Variable | Default | Description |
|---|---|---|
RENDITION_CODEC | h264 | h264 or hevc — base codec for every rendition |
MAX_RENDITION_HEIGHT | 0 | Cap the tallest rendition (0 = full 360p–4320p) — e.g. 2160 drops the niche 4320p rung |
HEVC_MIN_HEIGHT | 0 | Hybrid ladder: renditions at/above this height encode in HEVC while lower rungs keep the base codec (0 = disabled) — e.g. 2160 gives 4K+ HEVC + H.264 below |
FFMPEG_HWACCEL | auto | auto | vaapi | nvenc | disabled — hardware acceleration for the ladder encode. auto probes NVENC (/dev/nvidia0) first, then VA-API (/dev/dri/renderD128) |
FORCE_APT_FFMPEG (Docker build arg, not a runtime env var) — the static ffmpeg build (default, see Dockerfile's FFMPEG_VERSION) segfaults on virtually every -c copy remux on any CPU exposing AVX-512 (confirmed: Intel Xeon Cascadelake, Vultr) — a statically-linked-glibc IFUNC dispatch bug, not a code/content issue; -cpuflags doesn't fix it. Set FORCE_APT_FFMPEG=true as a build-time env var (Coolify forwards it as a real --build-arg) to use the apt-installed dynamically-linked ffmpeg package instead — same branch ENABLE_VAAPI/ENABLE_NVENC already use. Full root-cause writeup + runbook: docs/dev/BUG-transcoder-audio-remux-segfault.md.
The profile/level/CodecTag column shows the H.264 rendering. With RENDITION_CODEC=hevc (or HEVC_MIN_HEIGHT tiering) the encoder switches to libx265/hevc_vaapi, the profile becomes main (HEVC), and the CodecTag becomes the HEVC hvc1.1.6.L{levelIdc}.B0 form (vaapi.CodecTag) — the levelIdc is derived from the H.264 ladder's level tag (informational only, the encoders don't enforce it). Note: 4K+ HEVC rungs only play on HEVC-capable clients (Safari, modern Chrome/Edge with hardware decode); the ABR falls back to the highest H.264 rung on devices without HEVC support.
Single-pass encode — TranscodeLadder runs the whole ladder as one ffmpeg process (buildLadderArgs): the source is decoded once, a filter_complex split fans the decoded frames out to one scaled branch per rendition, and each branch writes its own HLS output — replacing the old N-decode × N-encode sequential passes, so total transcode time drops to roughly the longest rendition's encode. Settings are identical to the old per-rendition commands:
- CPU path:
libx264/libx265(RENDITION_CODEC=hevc) with CRF 23/28 + maxrate/bufsize (2× maxrate),-preset fast, profile/level,-tag:v hvc1for HEVC; every scale filter pinsformat=yuv420p—-profile:v main/highrejects 4:4:4 sources (e.g.testsrc, some screen recordings) - VA-API path (
FFMPEG_HWACCEL=vaapi, Intel/AMD):h264_vaapi/hevc_vaapiwith CBR (-rc_mode CBR+ maxrate/bufsize); decode+scale run on the GPU (internal/vaapi) - NVENC path (
FFMPEG_HWACCEL=nvenc, NVIDIA):h264_nvenc/hevc_nvencwith CBR (-rc cbr+ maxrate/bufsize) — encode-only: decode and scale still run on the CPU path above (nvenc accepts plainyuv420pframes directly, no hw-frame upload needed, seeinternal/nvenc) - 2s keyframes (
-force_key_frames expr:gte(t,n_forced*2)), HLS segments 6s, VOD playlist type - Dynamic timeout:
max(5min, duration×4s)for the whole ladder
Rich master playlist (CreateMasterPlaylist) — declares #EXT-X-VERSION:6 + #EXT-X-INDEPENDENT-SEGMENTS (every segment decodes standalone — guaranteed by the ladder's 2s forced keyframes) and enriches each #EXT-X-STREAM-INF with the source's FRAME-RATE (the ladder never resamples, from ffprobe's avg_frame_rate) and a measured AVERAGE-BANDWIDTH: the produced segments' total bits divided by the Media Playlist duration (RFC 8216), clamped to BANDWIDTH, falling back to the configured target when segments are missing.
Shared audio (EXT-X-MEDIA) — instead of re-encoding AAC into every rendition, the same pass emits one shared audio track to audio/index.m3u8, declared in master.m3u8 as an EXT-X-MEDIA AUDIO group (GROUP-ID="audio", DEFAULT=YES, AUTOSELECT=YES, LANGUAGE="und"). Renditions are video-only (-an) and mark AUDIO="audio" on their #EXT-X-STREAM-INF, with BANDWIDTH, AVERAGE-BANDWIDTH (audio's measured segments added) and CODECS including the audio track (mp4a.40.2). Settings come from AUDIO_PLAYBACK_BITRATE_KBPS (128k) / AUDIO_PLAYBACK_SAMPLE_RATE (48k) / AUDIO_PLAYBACK_CHANNELS (2). Sources without an audio track (probe HasAudio=false) skip the audio rendition entirely — silent assets still reach ready.
Other ladder outputs:
master.m3u8— written byCreateMasterPlaylist(VERSION 6,EXT-X-INDEPENDENT-SEGMENTS, per-variantFRAME-RATE+ measuredAVERAGE-BANDWIDTH, sharedEXT-X-MEDIAAUDIO group)download.mp4per rendition — fast-c copyremux of the rendition playlist plus the audio playlist when present (CreateDownloadableMp4),+faststartthumbnail.jpg— poster at 25% duration (ExtractPosterThumbnail)thumbnails/sprite_%03d.jpg— scrub-bar thumbs tiled at ~100 thumbs per file (5×20 grid of 160px-wide thumbs, 5s interval) in a single ffmpeg pass (thumbnails.Generate— one decode, aselectwindow +tileper file), so a hover preview downloads one small JPEG instead of one unbounded sheet (the old single sprite also hit libjpeg's 65500px dimension ceiling around 18h of video). All tiles in a multi-tile run share identical dimensions (the tile filter pads the short final tile), and a single-tile video is sized to exactly the rows it uses — both keep the player's single global sprite-size math correct.thumbnails.vttreferences each file assprite_%03d.jpg#xywh=...audio.m4a— the public-download AAC file, derived from the shared EXT-X-MEDIA audio rendition by stream copy (CreateDownloadableM4a,-c copyfromaudio/index.m3u8) — no second encode; it inherits the HLS track'sAUDIO_PLAYBACK_*settings exactly- AI audio —
ExtractAiAudiois a separate low-quality mono MP3 (AUDIO_AI_BITRATE_KBPS64k / 16k mono) for Whisper
Scaling & Hardware Adaptation
| Variable | Used by | Default | Description |
|---|---|---|---|
WORKER_CONCURRENCY | Transcoder (Go) | auto (CPU/RAM) | Concurrent transcode jobs — override for the auto formula |
FFMPEG_THREADS | Transcoder (Go) | auto | Threads per FFmpeg process (fed to the single-pass -threads) |
DB_POOL_SIZE | API, Transcoder | auto | Database connection pool size |
Transcoder sizing is cgroup-aware (internal/config): concurrency = min(availableMem/1.5GB, cpuCores/threadsPerJob) with threadsPerJob capped at 4, FFmpeg threads = cpuCores ÷ concurrency, pool = concurrency×2+2. It reads container memory limits (cgroup v2 → v1 → /proc/meminfo) instead of host totals, fixing the old Node worker's os.totalmem() overestimate in memory-limited containers.
Environment Variables
Defined in .env.example. apps/api-edge/apps/tus-edge env comes from wrangler.toml [vars] + wrangler secret (typed per-binding, no Zod). Dashboard and Player use Vite VITE_* build-time vars.
Key vars:
SHARED_AUTH_SECRET— base64 keyapps/api-edgeuses to sign (upload-token) andapps/tus-edgeuses to verify the TUS upload JWT. Generate withopenssl rand -base64 32(seescripts/init-shared-auth-secret.sh). Set as awrangler secreton both Workers.VITE_TUS_SERVER_URL— dashboard build-time base URL the TUS client appends/upload/videosto — theapps/tus-edgeWorker's origin (notVITE_API_BASE_URL/api-edge).VITE_PLAYER_BASE_URL— dashboard build-time URL of the player app (for embed link generation)QSTASH_TOKEN/QSTASH_CURRENT_SIGNING_KEY/QSTASH_NEXT_SIGNING_KEY— used for the transcode dispatch QStash publishes (/qstash/jobstoapps/transcoder). The analytics/backup/org-delete crons have been migrated to trigger.dev scheduled tasks and no longer use QStash. Get token + signing keys at https://console.upstash.com/qstash;QSTASH_URLoverrides the SDK base URL (local dev/tests point it at a stub).TRANSCODER_PUBLIC_URL(set onapps/api-edge) — QStash publish target for the transcoder's/qstash/jobsconsumer.WHISPER_WEBHOOK_SECRET— HMAC secret Modal signs its transcription callback with (X-Signature: sha256=<hmac>over the raw body). Only needed whenTRANSCRIPTION_PROVIDER=modal. Must be identical to thewhisper-webhook-secretModal Secret (seeinfra/modal-whisper/README.md);apps/api-edgeverifies it onPOST /v1/ai/whisper-callback. Generate withopenssl rand -hex 32.API_PUBLIC_URL— publicly reachable base URL ofapps/api-edge(only formodaltranscription). Thetranscode-orchestrationtask builds the Modalcallback_urlfrom it (<API_PUBLIC_URL>/v1/ai/whisper-callback); Modal calls back from its own cloud, solocalhostwon't work — use a tunnel (ngrok, Cloudflare Tunnel) or the deployedvapi.usestrum.appin dev.FLY_TRANSCODER_APP(set onapps/api-edge'swrangler.toml) — Fly app name ofapps/transcoder(strum-vod-transcoder). Used for diagnostics/queue-lag reporting (routes/health.ts,services/worker-status.ts) — no longer for waking it (that'stranscode-orchestration'swakeFlyApp, keyed on themachinestable'swakeStrategy/flyApp, not this env var).SELF_STOP_IDLE_SECONDS(optional, set onapps/transcoder) — how longinternal/selfstopwaits with zero active jobs before exiting so Fly stops the machine. Defaults to 120. No Fly API token needed — see "Worker architecture" above.MUX_DATA_ENV_KEY(optional, set onapps/api-edge) — platform-wide fallback Mux Data monitoring key, used when an org hasn't set its own via the dashboard Settings page (settings.mux_data_env_key). Per-org key wins (playback.service.ts). This is QoE/rebuffer/startup-time analytics only — not Mux Video; this project's HLS ladder is its own Go transcoder, unrelated to Mux's encoding product.
Mux Data (QoE analytics)
Real Mux Data integration via the mux-embed npm package (packages/player-ui, apps/player, apps/dashboard all declare it) — not a stub.
- Config —
settings.mux_data_env_keycolumn (packages/db/src/schema.ts), read/written viaGET/PATCH /v1/settings(apps/api-edge/src/services/settings.service.ts,handlers/settings.handler.ts). Falls back to theMUX_DATA_ENV_KEYbinding when no org-level key is set (services/playback.service.ts, exposed on the public playback payload asmuxDataEnvKey). - Player —
packages/player-ui/src/video-player.tsxcallsmux.monitor(videoEl, { data: { env_key, video_id, video_title, player_name, video_duration, ... } })on mount (hls.js instance passed in for HLS-level QoE),mux.destroyMonitor()on unmount. Wrapped in try/catch — analytics failures never break playback.EmbedPlayer.tsxthreadsmuxDataEnvKeydown from the playback payload. - Dashboard —
SettingsPage.tsxhas the key input;AnalyticsPage.tsxshows configured/not-configured status. - CLI —
apps/cli/src/commands/analytics.tsprints the configured key masked. - Not documented anywhere else before 2026-08-22 (no README/architecture.md mention) despite being fully wired end to end — this section is the first.
Agent skills
Issue tracker
Issues live in GitHub Issues (gh CLI). See docs/agents/issue-tracker.md.
Triage labels
Defaults: needs-triage, needs-info, ready-for-agent, ready-for-human, wontfix. See docs/agents/triage-labels.md.
Domain docs
Single-context: CONTEXT.md + docs/adr/ at repo root. See docs/agents/domain.md.
Trigger.dev agent skills
This project has Trigger.dev agent skills installed in .claude/skills/. Before writing or changing Trigger.dev code (background tasks, scheduled tasks, realtime, or chat.agent AI agents), load the most relevant skill: trigger-authoring-chat-agent, trigger-authoring-tasks, trigger-chat-agent-advanced, trigger-cost-savings, trigger-getting-started, trigger-realtime-and-frontend.