Running a Worker (VPS, Bare Metal, Your Own Machine)
This is the operational runbook for adding transcode capacity to production by running the worker binary on any machine you control — a VPS, a dedicated server, a home rig, even your own laptop for a burst of extra capacity. No Fly deploy, no code change, no downtime for the rest of the platform. This is the exact procedure used to bring the first non-Fly worker online in production (2026-08-16) — see root CLAUDE.md's "Worker unification" section for the design.
What you're running
A worker is one Go binary — apps/transcoder/cmd/transcoder — that:
- Never touches Postgres, Redis, or R2 credentials directly. Every job claim, progress update, file upload (via presigned URLs it's handed) and completion goes through
apps/api-edge's/v1/worker-agent/*HTTP broker, authenticated with a scoped machine token (mt_live_...). The one optional exception is anLISTEN/NOTIFYconnection string (DATABASE_URL) used only to wake instantly on new jobs — and that must be a least-privilegeworker_listenerrole (see below), never the app's owner credentials. - Pulls, never gets pushed to. It polls
GET /v1/worker-agent/jobs/claimevery 5s and claims whatever's next (FOR UPDATE SKIP LOCKEDserver-side — safe to run many workers against the same token with zero coordination). - Auto-detects its own hardware (CPU cores, RAM, cgroup limits if containerized) and sizes ffmpeg concurrency/thread count accordingly — no tuning needed to get started.
- Handles all 4 job types: full transcode ladder, audio extraction, stem mixing, and highlight-clip rendering.
- Ships its own tiny local web GUI (bound to
127.0.0.1only) so you can watch it work without SSHing in and tailing logs.
Because it never holds real credentials, a compromised or stolen worker machine cannot read/write your database or your R2 bucket — the worst case is someone burns compute running fake transcode jobs. That window is bounded by the token's lifetime: mint throwaway tokens with a ttlDays lease (see Step 1) so a stolen token dies on its own, and always revoke the token when you decommission a machine (see Revoking access). If you enable the optional LISTEN/NOTIFY wake path, the worker_listener connection is locked down to CONNECT-only, so a compromised worker still cannot read or write any data.
Requirements
- OS: Linux or macOS (Windows works too, cross-compiled, see Building from source).
- ffmpeg + ffprobe on
PATH. Not bundled — install via your package manager (apt install ffmpeg,brew install ffmpeg, etc.). This is the only hard runtime dependency. - yt-dlp on
PATH— optional, only needed if you process assets imported by URL instead of direct upload. - Go 1.23+ toolchain — only if building from source (recommended today, see below).
- Docker — only if you'd rather run the container image.
- Outbound HTTPS to your API's public URL. No inbound ports required — the worker only makes outbound polling requests. The local GUI binds to
127.0.0.1, not0.0.0.0, so it's not reachable from outside the machine by default either.
Step 1 — Mint a machine token
Machine tokens are a superadmin-only credential, minted once and reused across as many physical machines as you want — you don't register machines individually; a machines row is created automatically the first time a worker using that token heartbeats.
Dashboard: log in as a superadmin, go to /admin/machines — the "Workers" tab. It shows every registered machine (online status, which token, hardware, jobs done/failed, last heartbeat), lets you mint/revoke tokens, revoke individual machines, and hands you the ready-to-copy install command the moment you create a token.
API, if you'd rather script it — get a superadmin JWT (log in as the account that owns the first org, or any account promoted to superadmin), then:
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens \
-H "Authorization: Bearer $JWT" \
-H 'Content-Type: application/json' \
-d '{"name":"vps-fra1"}'{ "data": { "id": "...", "name": "vps-fra1", "token": "mt_live_...", "prefix": "mt_live_xxx" } }Copy the token value now — it's shown once. The DB only ever stores a hash of it (same pattern as org API keys), so there's no "reveal" later — losing it means minting a new one. The prefix is safe to keep around for identifying which token is which in logs/audits.
Ephemeral tokens (optional ttlDays)
By default a token has no expiry — it lives until you manually revoke it (see Revoking access). For throwaway/ephemeral workers — a laptop, a burst VPS you spin up and tear down, a friend's box — prefer a bounded lifetime so a leaked or stolen token stops working on its own even if nobody revokes it:
# Expires in 7 days — after that, all claim/heartbeat calls 401 and the
# worker is effectively dead until you mint a fresh token.
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens \
-H "Authorization: Bearer $JWT" \
-H 'Content-Type: application/json' \
-d '{"name":"laptop-local","ttlDays":7}'ttlDays is 1–365. The dashboard's Workers tab exposes the same option when creating a token and shows each token's status ("Ativo · expira …" / "Expirado") so you can spot a dead fleet at a glance. Expiry is enforced server-side in requireWorkerAuth (apps/api-edge/src/middleware/auth.ts) at the moment of every request, so an expired token can't claim, heartbeat, or upload even while it's still in the DB. The expiresAt timestamp is returned on create and included in GET /v1/admin/machine-tokens.
Step 2 — Get the binary
Option A — Build from source (recommended today)
This is the only verified working path in production right now. A prebuilt-binary download flow exists in the code (GET /install-worker.sh, GET /worker/download/:asset, backed by .github/workflows/release-worker.yml) but that workflow has never actually been triggered yet — no worker-v* tag has been pushed, so the download URL 404s today. See Publishing prebuilt binaries — it's a two-command fix, worth doing once you've confirmed the manual build works. Until then, build from source:
git clone <this repo> # or scp/rsync just apps/transcoder/ to the machine
cd strum-vod/apps/transcoder
go build -o strum-transcoder ./cmd/transcoderStatic Go binary, no runtime dependencies beyond ffmpeg/ffprobe on PATH. Cross-compile for another OS/arch from any machine with the Go toolchain:
GOOS=linux GOARCH=amd64 go build -o strum-transcoder-linux-amd64 ./cmd/transcoder
GOOS=linux GOARCH=arm64 go build -o strum-transcoder-linux-arm64 ./cmd/transcoder
GOOS=darwin GOARCH=arm64 go build -o strum-transcoder-darwin-arm64 ./cmd/transcoder
GOOS=windows GOARCH=amd64 go build -o strum-transcoder.exe ./cmd/transcoderCopy the resulting binary to the target machine (no Go toolchain needed there — it's fully static).
Option B — Docker
apps/transcoder/Dockerfile builds a self-contained image (static ffmpeg + yt-dlp baked in, so the host doesn't need them installed at all):
cd strum-vod
docker build -t strum-transcoder -f apps/transcoder/Dockerfile .
docker run -d --name strum-transcoder --restart unless-stopped \
-e STRUM_API_URL=https://vapi.usestrum.app \
-e STRUM_NODE_TOKEN=mt_live_... \
strum-transcoderPass --build-arg ENABLE_VAAPI=true to the docker build if the host has a usable /dev/dri VA-API device and you want hardware-accelerated encoding inside the container — then also add --device /dev/dri to docker run. For an NVIDIA GPU, pass --build-arg ENABLE_NVENC=true instead, and run with --gpus all -e NVIDIA_DRIVER_CAPABILITIES=compute,video,utility (needs nvidia-container-toolkit on the host — see Hardware acceleration for details). Otherwise the image ships a static (software-only) ffmpeg build, which is the safer default.
Option C — Prebuilt binary via install script (not yet available)
curl -fsSL https://vapi.usestrum.app/install-worker.sh \
| bash -s -- --token=mt_live_... --api-url=https://vapi.usestrum.appDetects OS/arch, checks for ffmpeg/ffprobe, downloads the binary, writes a config file, and prints the run command — the whole flow already works except the binary it tries to download doesn't exist in R2 yet. Skip this until you've done the one-time release setup in Publishing prebuilt binaries.
Step 3 — Run it
./strum-transcoder --api-url https://vapi.usestrum.app --token mt_live_...or via environment variables (equivalent, useful for systemd/Docker):
export STRUM_API_URL=https://vapi.usestrum.app
export STRUM_NODE_TOKEN=mt_live_...
./strum-transcoderOn startup it prints its detected hardware config and the local GUI URL, then starts polling for jobs immediately — no further setup needed.
[transcoder] hardware-adaptive config:
CPU cores: 8
Total RAM: 16.0 GB
Concurrency: 2 job(s)
FFmpeg threads: 4 per job
Hardware accel: false
[transcoder] local GUI: http://127.0.0.1:8791
[transcoder] connecting to https://vapi.usestrum.app...CLI flags / env vars reference
| Flag | Env var | Default | Description |
|---|---|---|---|
--api-url | STRUM_API_URL | — (required) | Your API's public base URL |
--token | STRUM_NODE_TOKEN | — (required) | Machine token from Step 1 (mt_live_...) |
--max-concurrency | STRUM_MAX_CONCURRENCY | 0 (auto) | Cap concurrent jobs; 0 = auto-detect from CPU/RAM |
--hwaccel | FFMPEG_HWACCEL | auto | auto | vaapi | nvenc | disabled — see below |
--verbose | — | false | Show ffmpeg's raw stdout/stderr instead of clean phase logging |
--no-gui | — | false | Disable the local web GUI entirely |
--gui-port | STRUM_NODE_GUI_PORT | 8791 | Local GUI port; 0 = auto-assign a free port |
--claim-poll-interval | STRUM_CLAIM_POLL_INTERVAL | 30 | Idle job-claim poll interval (seconds) — see below |
| — | STRUM_NODE_AGENT_INSTALL_DIR | ~/.strum-vod | Where the machine-id file, worker registry + logs live |
| — | MACHINE_WAKE_STRATEGY | none | fly_api only matters for Fly-hosted workers, irrelevant on a VPS/laptop — leave unset |
| — | SELF_STOP_IDLE_SECONDS | 120 | Scale-to-zero idle timeout — a no-op unless FLY_MACHINE_ID is set, so irrelevant here |
Throttle --max-concurrency on a shared/low-power machine if you don't want the worker using every core; leave it at 0 on a dedicated box.
Claim poll interval (STRUM_CLAIM_POLL_INTERVAL)
New jobs reach a worker instantly via Postgres LISTEN/NOTIFY, so this ticker is not the primary wake path — it's only the fallback that covers two NOTIFY gaps:
- a job dispatched while the worker's LISTEN connection is briefly down (a lost
NOTIFYis never redelivered), and - a lease-expired job reclaimed by
reclaim-stale-dispatch-jobsthat resets toqueuedwithout notifying.
Every idle concurrency slot polls once per interval, so the fallback volume is concurrency / interval requests per second against your API edge (Cloudflare Workers, metered per-request). The default 30 is a good balance; on an always-on dedicated machine (VPS/home rig) you can safely raise it to 120 or more to cut that request volume, since LISTEN/NOTIFY handles the common case with no latency penalty.
Optional: instant wake via Postgres LISTEN/NOTIFY (DATABASE_URL)
By default the worker only polls for jobs, which adds up to interval seconds of latency on an idle-to-busy transition. If you pass DATABASE_URL, the worker also opens a Postgres LISTEN connection and wakes the moment a NOTIFY job_queue fires — jobs start ~instantly with no polling gap.
This is the only time the worker holds a database credential, and it must be least-privilege. Do not reuse the app's neondb_owner connection string: a worker with neondb_owner could read/write your entire schema. A LISTEN needs almost no permissions — CONNECT on the database plus a NOTIFY/LISTEN capability requires no table privileges at all. Create a dedicated role:
-- Run as a privileged user (e.g. the DB owner) — once per database.
CREATE ROLE worker_listener LOGIN PASSWORD '<a-long-random-password>';
GRANT CONNECT ON DATABASE "yourdb" TO worker_listener;
-- LISTEN/NOTIFY need no table/column grants. Done.Then build a connection string for it, e.g.:
DATABASE_URL='postgres://worker_listener:<password>@<host>:5432/yourdb?sslmode=require'
LISTENis connection-scoped — it must be a persistent connection, not a request-per-job pool. The worker keeps one idleLISTENconnection open for the lifetime of the process. Grantingworker_listeneronlyCONNECTmeans even a fully compromised worker canLISTEN(and nothing else) — it can read no data, write nothing, and see none of your other schemas.
Already provisioned (this repo): prod role worker_listener (CONNECT-only) in scripts/.env.worker-listener, and a staging role with the same name but a distinct password in scripts/.env.worker-listener.staging — separate secrets per environment, so a leaked staging credential can't touch prod. Two helpers wrap all of this:
scripts/provision-worker-listener.sh <prod|staging> [--rotate]— creates the role, grantsCONNECT, verifiesLISTEN, and writes the env file.scripts/start-worker.sh [prod|staging|local] [--hwaccel ...]— sources the machine token (gitignoredscripts/.env.worker-token) + the listener env file and launches the node. For manual setups, source the env file yourself:
set -a; . scripts/.env.worker-listener; set +a
./apps/transcoder/strum-transcoder --hwaccel disabled(For a brand-new environment, run the CREATE ROLE/GRANT CONNECT above once, fill in the password, and save the URL in a gitignored env file.)
Without DATABASE_URL the worker degrades gracefully to pure polling (the default behavior above) — nothing else changes.
Running it permanently
A worker that dies when you close your terminal or reboots without restarting isn't useful in production. Pick one:
systemd (Linux, recommended for a VPS)
# /etc/systemd/system/strum-transcoder.service
[Unit]
Description=Strum VOD worker
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=strum
WorkingDirectory=/opt/strum-transcoder
Environment=STRUM_API_URL=https://vapi.usestrum.app
Environment=STRUM_NODE_TOKEN=mt_live_...
ExecStart=/opt/strum-transcoder/strum-transcoder
Restart=always
RestartSec=5
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable --now strum-transcoder
sudo journalctl -u strum-transcoder -f # tail logsPrefer a dedicated non-root user (useradd --system strum) — the binary never needs root, it only writes to its own install dir and talks HTTPS outbound.
Docker (--restart unless-stopped)
Already shown in Option B above — Docker's own restart policy covers reboots/crashes, no systemd unit needed on top.
tmux/screen (quick and dirty, dev/testing only)
tmux new -d -s transcoder './strum-transcoder --api-url ... --token ...'Fine for a quick test; doesn't survive a reboot, use systemd or Docker for anything you actually depend on.
Local GUI
Every worker serves a small dashboard at http://127.0.0.1:<gui-port> (default 8791) — broker connection state, detected hardware, active jobs with their current pipeline step, a running done/failed count, and a newest-first event log. A "Pause" button stops it claiming new work without killing the process (heartbeat keeps running, existing jobs finish); "Shut down" triggers the same graceful shutdown as Ctrl-C/SIGTERM. A "Cancel" button on each active job lets you terminate a single in-flight job immediately (marks it cancelled on the API side, re-queues the asset for retry if desired).
It's bound to 127.0.0.1 only and never receives the machine token — nothing sensitive crosses it. To view it on a headless VPS, SSH tunnel:
ssh -L 8791:127.0.0.1:8791 you@your-vps
# then open http://127.0.0.1:8791 locallyPrometheus metrics
The GUI also exposes a /metrics endpoint (Prometheus text format) at http://127.0.0.1:<gui-port>/metrics with:
- Machine info (CPU, RAM, GPU, OS/arch)
- Uptime, connection status, heartbeat errors
- Concurrency, paused state, active jobs count
- Per-job metrics (step, duration)
- Total jobs completed/failed counters
Add this as a Prometheus scrape target (e.g. via file_sd_configs or service discovery) to monitor your fleet.
Local worker registry (transcoder-ctl list)
Each worker auto-registers itself in ~/.strum-vod/workers/ on startup (writes a JSON file with machineId, guiUrl, pid, startedAt, hostname) and unregisters on clean shutdown. This lets you discover all workers running on the local machine without any API dependency:
# Build the CLI tool
cd strum-vod/apps/transcoder
go build -o transcoder-ctl ./cmd/transcoder-ctl
# List all local workers
./transcoder-ctl list
# MACHINE ID GUI URL PID STARTED HOSTNAME STATUS
# worker-abc123 http://127.0.0.1:8791 12345 2024-01-15T10:30:00Z mypc alive
# worker-def456 http://127.0.0.1:8792 12346 2024-01-15T10:31:00Z mypc dead (stale registry)The STATUS column checks if the PID is still alive — dead (stale registry) means the worker crashed hard (kill -9, power loss) and left its registry file behind; safe to ignore or clean up manually.
Other transcoder-ctl commands (work against a specific worker via --addr):
./transcoder-ctl status # Full status (hardware, connection, jobs)
./transcoder-ctl jobs # List active jobs
./transcoder-ctl pause # Pause job claiming
./transcoder-ctl resume # Resume job claiming
./transcoder-ctl cancel <job-id> # Cancel a specific running job
./transcoder-ctl shutdown # Graceful shutdown
./transcoder-ctl metrics # Prometheus metrics
./transcoder-ctl events # Recent events
./transcoder-ctl watch # Live dashboard (5s refresh)
./transcoder-ctl --addr http://192.168.1.100:8791 status # Remote workerHardware acceleration
--hwaccel auto (the default) probes NVENC first (/dev/nvidia0), then VA-API (/dev/dri), and uses whichever it finds. This can misfire: a machine with a device node present but non-functional (wrong driver, no permissions, a VM with a stub device) gets selected anyway, and the encode dies mid-job with a cryptic ffmpeg error instead of an obvious "no hardware accel" message. If you hit that, or you're unsure whether the machine's GPU setup actually works, pass --hwaccel disabled explicitly — safe, software-only (libx264/libx265), slower but always correct. Only reach for --hwaccel vaapi/--hwaccel nvenc (force one on) once you've confirmed the device genuinely works on that box.
NVENC (NVIDIA) — encode-only: decode/scale still run on the CPU, only h264_nvenc/hevc_nvenc do the encode, so it needs less GPU-side plumbing than VA-API. Requirements on the host:
- NVIDIA driver installed and
nvidia-smiworking. - ffmpeg built with
--enable-nvenc— Debian/Ubuntu's aptffmpegpackage already is; the johnvansickle static build used by Option A/B above is not. If you built from source (Option A), install ffmpeg from apt (apt install ffmpeg) or another NVENC-enabled build and make sure it's the one onPATHbefore the static binary, if any. - Docker (Option B): pass
--build-arg ENABLE_NVENC=true(switches the image to apt ffmpeg) and run the container with--gpus allplus-e NVIDIA_DRIVER_CAPABILITIES=compute,video,utility— thevideocapability is what exposeslibnvidia-encode.so; without it the encoder is listed but fails to open at runtime. Needs nvidia-container-toolkit installed on the host.
Scaling out
Run the same machine token on as many machines as you want — there's no per-token machine limit and no coordination needed between them (the claim query is FOR UPDATE SKIP LOCKED, so two workers can never grab the same job). Use one token per logical group if you want separate visibility in GET /v1/admin/machines (e.g. "office-rigs" vs "vps-fleet"), or one token for everything — it's purely a naming/reporting convenience, not a capacity limit.
Monitoring
The dashboard's /admin/machines page is the primary way to check this — online/offline status, active job count, and completed/failed counts update every 10s. Scripting it directly:
curl -s https://vapi.usestrum.app/v1/admin/machines \
-H "Authorization: Bearer $JWT" | jqReturns every registered machine: hardware, activeJobCount, jobsCompleted/jobsFailed, lastHeartbeatAt, and a computed online flag. A worker heartbeats every 5s; the server considers it offline after 15s without one (3 missed beats) — so online: false within seconds of you killing a worker is expected, not a bug.
Revoking access
Two levels, both superadmin-only, both one click on /admin/machines ("Revogar" on a token row or a machine row) — or scripted:
# Revoke ONE physical machine (e.g. decommissioning a specific VPS)
curl -X POST https://vapi.usestrum.app/v1/admin/machines/:machineId/revoke \
-H "Authorization: Bearer $JWT"
# Revoke the TOKEN itself — every machine using it stops being able to
# claim/heartbeat immediately, whether you know how many there are or not
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens/:tokenId/revoke \
-H "Authorization: Bearer $JWT"Revoking a token is the right move if it may have leaked (committed to a repo, pasted somewhere public, a machine running it was compromised) — do that first, then mint a fresh one and re-roll every machine that was using the old one.
Troubleshooting
| Symptom | Likely cause |
|---|---|
Worker logs connecting to ... then nothing, no jobs ever claimed | Token revoked, wrong --api-url, or the platform kill switch/tier gate — check GET /v1/admin/machines shows the machine registered at all (if it's not even listed, the token itself is bad) |
| ffmpeg crashes mid-job with a filter-graph error, no clear message | VA-API misdetection — restart with --hwaccel disabled, see Hardware acceleration |
Cannot load libnvidia-encode.so.1 or NVENC "no capable devices found" mid-job | Container missing NVIDIA_DRIVER_CAPABILITIES=compute,video,utility (needs video) or --gpus all, or the host driver/nvidia-container-toolkit isn't installed — see Hardware acceleration |
strum-transcoder: --api-url and --token are required | Neither the flags nor STRUM_API_URL/STRUM_NODE_TOKEN env vars are set |
| GUI unreachable from another machine | Bound to 127.0.0.1 by design — SSH tunnel (see Local GUI), or accept it's local-only |
GET /worker/download/... returns 404 during the install script | Expected until you've run the release workflow once — see Option C / Publishing prebuilt binaries. Build from source instead. |
Machine shows online: false right after you know it's running | Heartbeat is every 5s / TTL is 15s — wait ~15s after startup before concluding it's actually down |
Publishing prebuilt binaries (optional)
If you want the curl \| bash one-liner to actually work (useful for a fleet you're growing regularly, or handing setup off to someone who shouldn't need a Go toolchain), the pipeline already exists — it's just never been triggered:
git tag worker-v1.0.0
git push origin worker-v1.0.0This fires .github/workflows/release-worker.yml, which cross-compiles cmd/transcoder for linux/darwin × amd64/arm64 plus windows/amd64, and uploads each to your R2 bucket at both releases/worker/latest/strum-vod-worker-<os>-<arch> (what /worker/download/:asset always redirects to) and releases/worker/worker-v1.0.0/strum-vod-worker-<os>-<arch> (pinned, for rollback). Requires the S3_ENDPOINT/S3_BUCKET/S3_ACCESS_KEY_ID/S3_SECRET_ACCESS_KEY GitHub Actions repo secrets to already be set (same R2 credentials the app itself uses; S3_PUBLIC_BASE_URL is optional, only used to print direct URLs in the run's job summary). Once that's run once, Option C above starts working, machine token and all — the install script never needed changes, only the binary needed to actually exist at that URL. You can also trigger it manually from the Actions tab (workflow_dispatch) against an existing tag, no new push required.
The pre-unification customer "Nodes" and admin "Platform Workers" fleets (nk_/pw_ tokens, cmd/node-agent, release-node-agent.yml) have been fully deleted — the unified fleet this guide describes is the only worker system left.