Skip to content

Running a Worker (VPS, Bare Metal, Your Own Machine) ​

This is the operational runbook for adding transcode capacity to production by running the worker binary on any machine you control — a VPS, a dedicated server, a home rig, even your own laptop for a burst of extra capacity. No Fly deploy, no code change, no downtime for the rest of the platform. This is the exact procedure used to bring the first non-Fly worker online in production (2026-08-16) — see root CLAUDE.md's "Worker unification" section for the design.

What you're running ​

A worker is one Go binary — apps/transcoder/cmd/transcoder — that:

  • Never touches Postgres, Redis, or R2 credentials directly. Every job claim, progress update, file upload (via presigned URLs it's handed) and completion goes through apps/api-edge's /v1/worker-agent/* HTTP broker, authenticated with a scoped machine token (mt_live_...). The one optional exception is an LISTEN/NOTIFY connection string (DATABASE_URL) used only to wake instantly on new jobs — and that must be a least-privilege worker_listener role (see below), never the app's owner credentials.
  • Pulls, never gets pushed to. It polls GET /v1/worker-agent/jobs/claim every 5s and claims whatever's next (FOR UPDATE SKIP LOCKED server-side — safe to run many workers against the same token with zero coordination).
  • Auto-detects its own hardware (CPU cores, RAM, cgroup limits if containerized) and sizes ffmpeg concurrency/thread count accordingly — no tuning needed to get started.
  • Handles all 4 job types: full transcode ladder, audio extraction, stem mixing, and highlight-clip rendering.
  • Ships its own tiny local web GUI (bound to 127.0.0.1 only) so you can watch it work without SSHing in and tailing logs.

Because it never holds real credentials, a compromised or stolen worker machine cannot read/write your database or your R2 bucket — the worst case is someone burns compute running fake transcode jobs. That window is bounded by the token's lifetime: mint throwaway tokens with a ttlDays lease (see Step 1) so a stolen token dies on its own, and always revoke the token when you decommission a machine (see Revoking access). If you enable the optional LISTEN/NOTIFY wake path, the worker_listener connection is locked down to CONNECT-only, so a compromised worker still cannot read or write any data.

Requirements ​

  • OS: Linux or macOS (Windows works too, cross-compiled, see Building from source).
  • ffmpeg + ffprobe on PATH. Not bundled — install via your package manager (apt install ffmpeg, brew install ffmpeg, etc.). This is the only hard runtime dependency.
  • yt-dlp on PATH — optional, only needed if you process assets imported by URL instead of direct upload.
  • Go 1.23+ toolchain — only if building from source (recommended today, see below).
  • Docker — only if you'd rather run the container image.
  • Outbound HTTPS to your API's public URL. No inbound ports required — the worker only makes outbound polling requests. The local GUI binds to 127.0.0.1, not 0.0.0.0, so it's not reachable from outside the machine by default either.

Step 1 — Mint a machine token ​

Machine tokens are a superadmin-only credential, minted once and reused across as many physical machines as you want — you don't register machines individually; a machines row is created automatically the first time a worker using that token heartbeats.

Dashboard: log in as a superadmin, go to /admin/machines — the "Workers" tab. It shows every registered machine (online status, which token, hardware, jobs done/failed, last heartbeat), lets you mint/revoke tokens, revoke individual machines, and hands you the ready-to-copy install command the moment you create a token.

API, if you'd rather script it — get a superadmin JWT (log in as the account that owns the first org, or any account promoted to superadmin), then:

bash
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens \
  -H "Authorization: Bearer $JWT" \
  -H 'Content-Type: application/json' \
  -d '{"name":"vps-fra1"}'
json
{ "data": { "id": "...", "name": "vps-fra1", "token": "mt_live_...", "prefix": "mt_live_xxx" } }

Copy the token value now — it's shown once. The DB only ever stores a hash of it (same pattern as org API keys), so there's no "reveal" later — losing it means minting a new one. The prefix is safe to keep around for identifying which token is which in logs/audits.

Ephemeral tokens (optional ttlDays) ​

By default a token has no expiry — it lives until you manually revoke it (see Revoking access). For throwaway/ephemeral workers — a laptop, a burst VPS you spin up and tear down, a friend's box — prefer a bounded lifetime so a leaked or stolen token stops working on its own even if nobody revokes it:

bash
# Expires in 7 days — after that, all claim/heartbeat calls 401 and the
# worker is effectively dead until you mint a fresh token.
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens \
  -H "Authorization: Bearer $JWT" \
  -H 'Content-Type: application/json' \
  -d '{"name":"laptop-local","ttlDays":7}'

ttlDays is 1–365. The dashboard's Workers tab exposes the same option when creating a token and shows each token's status ("Ativo · expira …" / "Expirado") so you can spot a dead fleet at a glance. Expiry is enforced server-side in requireWorkerAuth (apps/api-edge/src/middleware/auth.ts) at the moment of every request, so an expired token can't claim, heartbeat, or upload even while it's still in the DB. The expiresAt timestamp is returned on create and included in GET /v1/admin/machine-tokens.

Step 2 — Get the binary ​

This is the only verified working path in production right now. A prebuilt-binary download flow exists in the code (GET /install-worker.sh, GET /worker/download/:asset, backed by .github/workflows/release-worker.yml) but that workflow has never actually been triggered yet — no worker-v* tag has been pushed, so the download URL 404s today. See Publishing prebuilt binaries — it's a two-command fix, worth doing once you've confirmed the manual build works. Until then, build from source:

bash
git clone <this repo>          # or scp/rsync just apps/transcoder/ to the machine
cd strum-vod/apps/transcoder
go build -o strum-transcoder ./cmd/transcoder

Static Go binary, no runtime dependencies beyond ffmpeg/ffprobe on PATH. Cross-compile for another OS/arch from any machine with the Go toolchain:

bash
GOOS=linux   GOARCH=amd64 go build -o strum-transcoder-linux-amd64   ./cmd/transcoder
GOOS=linux   GOARCH=arm64 go build -o strum-transcoder-linux-arm64   ./cmd/transcoder
GOOS=darwin  GOARCH=arm64 go build -o strum-transcoder-darwin-arm64  ./cmd/transcoder
GOOS=windows GOARCH=amd64 go build -o strum-transcoder.exe           ./cmd/transcoder

Copy the resulting binary to the target machine (no Go toolchain needed there — it's fully static).

Option B — Docker ​

apps/transcoder/Dockerfile builds a self-contained image (static ffmpeg + yt-dlp baked in, so the host doesn't need them installed at all):

bash
cd strum-vod
docker build -t strum-transcoder -f apps/transcoder/Dockerfile .

docker run -d --name strum-transcoder --restart unless-stopped \
  -e STRUM_API_URL=https://vapi.usestrum.app \
  -e STRUM_NODE_TOKEN=mt_live_... \
  strum-transcoder

Pass --build-arg ENABLE_VAAPI=true to the docker build if the host has a usable /dev/dri VA-API device and you want hardware-accelerated encoding inside the container — then also add --device /dev/dri to docker run. For an NVIDIA GPU, pass --build-arg ENABLE_NVENC=true instead, and run with --gpus all -e NVIDIA_DRIVER_CAPABILITIES=compute,video,utility (needs nvidia-container-toolkit on the host — see Hardware acceleration for details). Otherwise the image ships a static (software-only) ffmpeg build, which is the safer default.

Option C — Prebuilt binary via install script (not yet available) ​

bash
curl -fsSL https://vapi.usestrum.app/install-worker.sh \
  | bash -s -- --token=mt_live_... --api-url=https://vapi.usestrum.app

Detects OS/arch, checks for ffmpeg/ffprobe, downloads the binary, writes a config file, and prints the run command — the whole flow already works except the binary it tries to download doesn't exist in R2 yet. Skip this until you've done the one-time release setup in Publishing prebuilt binaries.

Step 3 — Run it ​

bash
./strum-transcoder --api-url https://vapi.usestrum.app --token mt_live_...

or via environment variables (equivalent, useful for systemd/Docker):

bash
export STRUM_API_URL=https://vapi.usestrum.app
export STRUM_NODE_TOKEN=mt_live_...
./strum-transcoder

On startup it prints its detected hardware config and the local GUI URL, then starts polling for jobs immediately — no further setup needed.

[transcoder] hardware-adaptive config:
  CPU cores:      8
  Total RAM:      16.0 GB
  Concurrency:    2 job(s)
  FFmpeg threads: 4 per job
  Hardware accel: false
[transcoder] local GUI: http://127.0.0.1:8791
[transcoder] connecting to https://vapi.usestrum.app...

CLI flags / env vars reference ​

FlagEnv varDefaultDescription
--api-urlSTRUM_API_URL— (required)Your API's public base URL
--tokenSTRUM_NODE_TOKEN— (required)Machine token from Step 1 (mt_live_...)
--max-concurrencySTRUM_MAX_CONCURRENCY0 (auto)Cap concurrent jobs; 0 = auto-detect from CPU/RAM
--hwaccelFFMPEG_HWACCELautoauto | vaapi | nvenc | disabled — see below
--verbose—falseShow ffmpeg's raw stdout/stderr instead of clean phase logging
--no-gui—falseDisable the local web GUI entirely
--gui-portSTRUM_NODE_GUI_PORT8791Local GUI port; 0 = auto-assign a free port
--claim-poll-intervalSTRUM_CLAIM_POLL_INTERVAL30Idle job-claim poll interval (seconds) — see below
—STRUM_NODE_AGENT_INSTALL_DIR~/.strum-vodWhere the machine-id file, worker registry + logs live
—MACHINE_WAKE_STRATEGYnonefly_api only matters for Fly-hosted workers, irrelevant on a VPS/laptop — leave unset
—SELF_STOP_IDLE_SECONDS120Scale-to-zero idle timeout — a no-op unless FLY_MACHINE_ID is set, so irrelevant here

Throttle --max-concurrency on a shared/low-power machine if you don't want the worker using every core; leave it at 0 on a dedicated box.

Claim poll interval (STRUM_CLAIM_POLL_INTERVAL) ​

New jobs reach a worker instantly via Postgres LISTEN/NOTIFY, so this ticker is not the primary wake path — it's only the fallback that covers two NOTIFY gaps:

  1. a job dispatched while the worker's LISTEN connection is briefly down (a lost NOTIFY is never redelivered), and
  2. a lease-expired job reclaimed by reclaim-stale-dispatch-jobs that resets to queued without notifying.

Every idle concurrency slot polls once per interval, so the fallback volume is concurrency / interval requests per second against your API edge (Cloudflare Workers, metered per-request). The default 30 is a good balance; on an always-on dedicated machine (VPS/home rig) you can safely raise it to 120 or more to cut that request volume, since LISTEN/NOTIFY handles the common case with no latency penalty.

Optional: instant wake via Postgres LISTEN/NOTIFY (DATABASE_URL) ​

By default the worker only polls for jobs, which adds up to interval seconds of latency on an idle-to-busy transition. If you pass DATABASE_URL, the worker also opens a Postgres LISTEN connection and wakes the moment a NOTIFY job_queue fires — jobs start ~instantly with no polling gap.

This is the only time the worker holds a database credential, and it must be least-privilege. Do not reuse the app's neondb_owner connection string: a worker with neondb_owner could read/write your entire schema. A LISTEN needs almost no permissions — CONNECT on the database plus a NOTIFY/LISTEN capability requires no table privileges at all. Create a dedicated role:

sql
-- Run as a privileged user (e.g. the DB owner) — once per database.
CREATE ROLE worker_listener LOGIN PASSWORD '<a-long-random-password>';
GRANT CONNECT ON DATABASE "yourdb" TO worker_listener;
-- LISTEN/NOTIFY need no table/column grants. Done.

Then build a connection string for it, e.g.:

DATABASE_URL='postgres://worker_listener:<password>@<host>:5432/yourdb?sslmode=require'

LISTEN is connection-scoped — it must be a persistent connection, not a request-per-job pool. The worker keeps one idle LISTEN connection open for the lifetime of the process. Granting worker_listener only CONNECT means even a fully compromised worker can LISTEN (and nothing else) — it can read no data, write nothing, and see none of your other schemas.

Already provisioned (this repo): prod role worker_listener (CONNECT-only) in scripts/.env.worker-listener, and a staging role with the same name but a distinct password in scripts/.env.worker-listener.staging — separate secrets per environment, so a leaked staging credential can't touch prod. Two helpers wrap all of this:

  • scripts/provision-worker-listener.sh <prod|staging> [--rotate] — creates the role, grants CONNECT, verifies LISTEN, and writes the env file.
  • scripts/start-worker.sh [prod|staging|local] [--hwaccel ...] — sources the machine token (gitignored scripts/.env.worker-token) + the listener env file and launches the node. For manual setups, source the env file yourself:
bash
set -a; . scripts/.env.worker-listener; set +a
./apps/transcoder/strum-transcoder --hwaccel disabled

(For a brand-new environment, run the CREATE ROLE/GRANT CONNECT above once, fill in the password, and save the URL in a gitignored env file.)

Without DATABASE_URL the worker degrades gracefully to pure polling (the default behavior above) — nothing else changes.

Running it permanently ​

A worker that dies when you close your terminal or reboots without restarting isn't useful in production. Pick one:

ini
# /etc/systemd/system/strum-transcoder.service
[Unit]
Description=Strum VOD worker
After=network-online.target
Wants=network-online.target

[Service]
Type=simple
User=strum
WorkingDirectory=/opt/strum-transcoder
Environment=STRUM_API_URL=https://vapi.usestrum.app
Environment=STRUM_NODE_TOKEN=mt_live_...
ExecStart=/opt/strum-transcoder/strum-transcoder
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target
bash
sudo systemctl daemon-reload
sudo systemctl enable --now strum-transcoder
sudo journalctl -u strum-transcoder -f   # tail logs

Prefer a dedicated non-root user (useradd --system strum) — the binary never needs root, it only writes to its own install dir and talks HTTPS outbound.

Docker (--restart unless-stopped) ​

Already shown in Option B above — Docker's own restart policy covers reboots/crashes, no systemd unit needed on top.

tmux/screen (quick and dirty, dev/testing only) ​

bash
tmux new -d -s transcoder './strum-transcoder --api-url ... --token ...'

Fine for a quick test; doesn't survive a reboot, use systemd or Docker for anything you actually depend on.

Local GUI ​

Every worker serves a small dashboard at http://127.0.0.1:<gui-port> (default 8791) — broker connection state, detected hardware, active jobs with their current pipeline step, a running done/failed count, and a newest-first event log. A "Pause" button stops it claiming new work without killing the process (heartbeat keeps running, existing jobs finish); "Shut down" triggers the same graceful shutdown as Ctrl-C/SIGTERM. A "Cancel" button on each active job lets you terminate a single in-flight job immediately (marks it cancelled on the API side, re-queues the asset for retry if desired).

It's bound to 127.0.0.1 only and never receives the machine token — nothing sensitive crosses it. To view it on a headless VPS, SSH tunnel:

bash
ssh -L 8791:127.0.0.1:8791 you@your-vps
# then open http://127.0.0.1:8791 locally

Prometheus metrics ​

The GUI also exposes a /metrics endpoint (Prometheus text format) at http://127.0.0.1:<gui-port>/metrics with:

  • Machine info (CPU, RAM, GPU, OS/arch)
  • Uptime, connection status, heartbeat errors
  • Concurrency, paused state, active jobs count
  • Per-job metrics (step, duration)
  • Total jobs completed/failed counters

Add this as a Prometheus scrape target (e.g. via file_sd_configs or service discovery) to monitor your fleet.

Local worker registry (transcoder-ctl list) ​

Each worker auto-registers itself in ~/.strum-vod/workers/ on startup (writes a JSON file with machineId, guiUrl, pid, startedAt, hostname) and unregisters on clean shutdown. This lets you discover all workers running on the local machine without any API dependency:

bash
# Build the CLI tool
cd strum-vod/apps/transcoder
go build -o transcoder-ctl ./cmd/transcoder-ctl

# List all local workers
./transcoder-ctl list
# MACHINE ID      GUI URL                PID    STARTED                 HOSTNAME  STATUS
# worker-abc123   http://127.0.0.1:8791  12345  2024-01-15T10:30:00Z    mypc      alive
# worker-def456   http://127.0.0.1:8792  12346  2024-01-15T10:31:00Z    mypc      dead (stale registry)

The STATUS column checks if the PID is still alive — dead (stale registry) means the worker crashed hard (kill -9, power loss) and left its registry file behind; safe to ignore or clean up manually.

Other transcoder-ctl commands (work against a specific worker via --addr):

bash
./transcoder-ctl status            # Full status (hardware, connection, jobs)
./transcoder-ctl jobs              # List active jobs
./transcoder-ctl pause             # Pause job claiming
./transcoder-ctl resume            # Resume job claiming
./transcoder-ctl cancel <job-id>   # Cancel a specific running job
./transcoder-ctl shutdown          # Graceful shutdown
./transcoder-ctl metrics           # Prometheus metrics
./transcoder-ctl events            # Recent events
./transcoder-ctl watch             # Live dashboard (5s refresh)
./transcoder-ctl --addr http://192.168.1.100:8791 status  # Remote worker

Hardware acceleration ​

--hwaccel auto (the default) probes NVENC first (/dev/nvidia0), then VA-API (/dev/dri), and uses whichever it finds. This can misfire: a machine with a device node present but non-functional (wrong driver, no permissions, a VM with a stub device) gets selected anyway, and the encode dies mid-job with a cryptic ffmpeg error instead of an obvious "no hardware accel" message. If you hit that, or you're unsure whether the machine's GPU setup actually works, pass --hwaccel disabled explicitly — safe, software-only (libx264/libx265), slower but always correct. Only reach for --hwaccel vaapi/--hwaccel nvenc (force one on) once you've confirmed the device genuinely works on that box.

NVENC (NVIDIA) — encode-only: decode/scale still run on the CPU, only h264_nvenc/hevc_nvenc do the encode, so it needs less GPU-side plumbing than VA-API. Requirements on the host:

  • NVIDIA driver installed and nvidia-smi working.
  • ffmpeg built with --enable-nvenc — Debian/Ubuntu's apt ffmpeg package already is; the johnvansickle static build used by Option A/B above is not. If you built from source (Option A), install ffmpeg from apt (apt install ffmpeg) or another NVENC-enabled build and make sure it's the one on PATH before the static binary, if any.
  • Docker (Option B): pass --build-arg ENABLE_NVENC=true (switches the image to apt ffmpeg) and run the container with --gpus all plus -e NVIDIA_DRIVER_CAPABILITIES=compute,video,utility — the video capability is what exposes libnvidia-encode.so; without it the encoder is listed but fails to open at runtime. Needs nvidia-container-toolkit installed on the host.

Scaling out ​

Run the same machine token on as many machines as you want — there's no per-token machine limit and no coordination needed between them (the claim query is FOR UPDATE SKIP LOCKED, so two workers can never grab the same job). Use one token per logical group if you want separate visibility in GET /v1/admin/machines (e.g. "office-rigs" vs "vps-fleet"), or one token for everything — it's purely a naming/reporting convenience, not a capacity limit.

Monitoring ​

The dashboard's /admin/machines page is the primary way to check this — online/offline status, active job count, and completed/failed counts update every 10s. Scripting it directly:

bash
curl -s https://vapi.usestrum.app/v1/admin/machines \
  -H "Authorization: Bearer $JWT" | jq

Returns every registered machine: hardware, activeJobCount, jobsCompleted/jobsFailed, lastHeartbeatAt, and a computed online flag. A worker heartbeats every 5s; the server considers it offline after 15s without one (3 missed beats) — so online: false within seconds of you killing a worker is expected, not a bug.

Revoking access ​

Two levels, both superadmin-only, both one click on /admin/machines ("Revogar" on a token row or a machine row) — or scripted:

bash
# Revoke ONE physical machine (e.g. decommissioning a specific VPS)
curl -X POST https://vapi.usestrum.app/v1/admin/machines/:machineId/revoke \
  -H "Authorization: Bearer $JWT"

# Revoke the TOKEN itself — every machine using it stops being able to
# claim/heartbeat immediately, whether you know how many there are or not
curl -X POST https://vapi.usestrum.app/v1/admin/machine-tokens/:tokenId/revoke \
  -H "Authorization: Bearer $JWT"

Revoking a token is the right move if it may have leaked (committed to a repo, pasted somewhere public, a machine running it was compromised) — do that first, then mint a fresh one and re-roll every machine that was using the old one.

Troubleshooting ​

SymptomLikely cause
Worker logs connecting to ... then nothing, no jobs ever claimedToken revoked, wrong --api-url, or the platform kill switch/tier gate — check GET /v1/admin/machines shows the machine registered at all (if it's not even listed, the token itself is bad)
ffmpeg crashes mid-job with a filter-graph error, no clear messageVA-API misdetection — restart with --hwaccel disabled, see Hardware acceleration
Cannot load libnvidia-encode.so.1 or NVENC "no capable devices found" mid-jobContainer missing NVIDIA_DRIVER_CAPABILITIES=compute,video,utility (needs video) or --gpus all, or the host driver/nvidia-container-toolkit isn't installed — see Hardware acceleration
strum-transcoder: --api-url and --token are requiredNeither the flags nor STRUM_API_URL/STRUM_NODE_TOKEN env vars are set
GUI unreachable from another machineBound to 127.0.0.1 by design — SSH tunnel (see Local GUI), or accept it's local-only
GET /worker/download/... returns 404 during the install scriptExpected until you've run the release workflow once — see Option C / Publishing prebuilt binaries. Build from source instead.
Machine shows online: false right after you know it's runningHeartbeat is every 5s / TTL is 15s — wait ~15s after startup before concluding it's actually down

Publishing prebuilt binaries (optional) ​

If you want the curl \| bash one-liner to actually work (useful for a fleet you're growing regularly, or handing setup off to someone who shouldn't need a Go toolchain), the pipeline already exists — it's just never been triggered:

bash
git tag worker-v1.0.0
git push origin worker-v1.0.0

This fires .github/workflows/release-worker.yml, which cross-compiles cmd/transcoder for linux/darwin × amd64/arm64 plus windows/amd64, and uploads each to your R2 bucket at both releases/worker/latest/strum-vod-worker-<os>-<arch> (what /worker/download/:asset always redirects to) and releases/worker/worker-v1.0.0/strum-vod-worker-<os>-<arch> (pinned, for rollback). Requires the S3_ENDPOINT/S3_BUCKET/S3_ACCESS_KEY_ID/S3_SECRET_ACCESS_KEY GitHub Actions repo secrets to already be set (same R2 credentials the app itself uses; S3_PUBLIC_BASE_URL is optional, only used to print direct URLs in the run's job summary). Once that's run once, Option C above starts working, machine token and all — the install script never needed changes, only the binary needed to actually exist at that URL. You can also trigger it manually from the Actions tab (workflow_dispatch) against an existing tag, no new push required.

The pre-unification customer "Nodes" and admin "Platform Workers" fleets (nk_/pw_ tokens, cmd/node-agent, release-node-agent.yml) have been fully deleted — the unified fleet this guide describes is the only worker system left.

STRUM Proprietary License — © 2026 Strum. All rights reserved.