17 KiB
id, title, created, depends_on, est_files
| id | title | created | depends_on | est_files | ||
|---|---|---|---|---|---|---|
| 007-ytdlp-worker-045800 | Keep one long-lived yt-dlp worker process instead of spawning per call | 2026-09-29 |
|
5 |
007 — Long-lived yt-dlp worker pool
Objective
Every yt-dlp call pays Python start-up + extractor import. Measured 2026-09-29 with yt-dlp 2026.08.19 (same host, same network):
| call | per-call spawn | warm worker |
|---|---|---|
ytsearch5 --flat-playlist |
2.49 s | 1.32 s |
-J <video> |
3.32 s | 2.29 s |
After this plan, the "read-only" yt-dlp calls (search fallback, channel, playlist
expand, -J stream resolve) go to a pool of 2 long-lived Python workers that import
yt_dlp once. Downloads (which write files and use an AbortSignal) keep spawning.
Any pool infrastructure problem falls back to spawning transparently. A bot-check answer
recycles the workers and retries that call with a fresh spawn; each worker is replaced after
100 requests. YTDLP_WORKER=0 disables the pool.
Dry-run notes (2026-09-29): the worker calls yt-dlp's own CLI entry (yt_dlp._real_main) — a
first version using parse_options() + YoutubeDL directly behaved differently from the CLI.
During testing the container's IP got bot-checked by YouTube for plain CLI calls too, so the
pooled -J path could only be measured before that (2.29 s vs 3.32 s). Measure on prod with
plan 001's [ytdlp] pooled … log lines and keep YTDLP_WORKER=0 as the escape hatch.
Context the executor must NOT rediscover
- In Docker, yt-dlp is the release zipapp at
/usr/local/bin/yt-dlp; Python can importyt_dlpfrom it by putting that path onsys.path.python3is installed in the image (Dockerfile apt line)./etc/yt-dlp.conf(the--js-runtimes bun:…line) is honoured becauseyt_dlp.parse_options()reads config files like the CLI does. - The Dockerfile copies
server/to/app, so a file atserver/ytdlp-worker.pyships automatically. server/server.js:143—function runYtdlp(args, { signal } = {})spawns yt-dlp (after plan 001 it also logs[ytdlp] <kind> <ms>ms ok|fail).server/server.js:189—async function runYtdlpResilient(args, opts = {})callsrunYtdlp(withCookies(args), opts)and, on a bot check, retries withrunYtdlp(withCookies(['--extractor-args', …, ...args]), opts).optsis passed through.- Read-only call sites to mark
{ pooled: true }(line numbers from before plans 004/006; search by the text):/api/searchfallback:runYtdlpResilient([\ytsearch${SEARCH_LIMIT}:${q}`, …])`/api/channel:runYtdlpResilient([…'--playlist-end', String(CHANNEL_LIMIT),resolveStreamsUncached:runYtdlpResilient(['-J', '--no-warnings', \https://www.youtube.com/watch?v=${videoId}`])`/api/playlist/expand(~line 1712):const out = await runYtdlpResilient([Do NOT mark the calls at ~861, ~932, ~988 (downloads) orserver/notes.js:362.
const YTDLP = process.env.YTDLP_PATH || 'yt-dlp';near the top of server.js.- Expected noise on a machine WITHOUT yt-dlp (local dev before
npm run setup): threeModuleNotFoundError: No module named 'yt_dlp'lines, then[ytdlp-pool] disabled after 3 failed starts …— the server keeps working via spawn. Not a bug.
Steps
- Create
server/ytdlp-worker.pywith exactly:#!/usr/bin/env python3 """ytdlp-worker.py - one long-lived yt-dlp process answering many requests. Spawning yt-dlp per call pays Python start-up + extractor import every time. This worker imports yt_dlp ONCE and runs each request's argv in-process through the CLI's own entry point (yt_dlp._real_main), so config files such as /etc/yt-dlp.conf and everything else the command line sets up apply. Protocol (newline-delimited JSON): stdin : {"id": <int>, "args": [<yt-dlp argv>...]} stdout: {"id": <int>, "code": <int>, "out": "<captured stdout>", "err": "<stderr tail>"} Only for calls whose result is printed to stdout (-J, --dump-json, --version). Usage: python3 ytdlp-worker.py <path-to-yt-dlp-zipapp-or-empty> """ import io, json, sys, contextlib if len(sys.argv) > 1 and sys.argv[1]: sys.path.insert(0, sys.argv[1]) # the yt-dlp release binary is a zipapp import yt_dlp # noqa: E402 proto_out = sys.stdout sys.stdout = io.StringIO() # nothing may leak onto the protocol pipe def run(args): out, err = io.StringIO(), io.StringIO() code = 0 with contextlib.redirect_stdout(out), contextlib.redirect_stderr(err): try: # The CLI's own entry point (not parse_options + YoutubeDL): it also # sets up what the command line does (JS challenge solving, plugins, # post-processing defaults). A bare YoutubeDL got "Sign in to # confirm you're not a bot" where the CLI succeeded. ret = yt_dlp._real_main(args) code = ret[0] if isinstance(ret, tuple) else (ret or 0) except SystemExit as e: code = e.code if isinstance(e.code, int) else 1 except Exception as e: # report, never die err.write(f"ERROR: {e}\n") code = 1 return code, out.getvalue(), err.getvalue()[-8000:] for line in sys.stdin: line = line.strip() if not line: continue try: req = json.loads(line) code, out, err = run([str(a) for a in req.get("args", [])]) resp = {"id": req.get("id"), "code": code, "out": out, "err": err} except Exception as e: resp = {"id": None, "code": 1, "out": "", "err": f"worker: {e}"} proto_out.write(json.dumps(resp) + "\n") proto_out.flush() - Create
server/ytdlp-pool.jswith exactly:/* ytdlp-pool.js — a small pool of long-lived yt-dlp workers (ytdlp-worker.py). * * run(args) resolves stdout exactly like a spawned `yt-dlp <args>` would, or * rejects with the captured stderr. Each worker handles one request at a time; * extra requests queue. A worker that exits or exceeds the timeout is killed * and replaced. Errors carrying `poolInfra: true` mean "the pool could not run * this" — the caller falls back to a normal per-call spawn. Three workers in a * row dying before answering anything (no python3, no importable yt_dlp) * disables the pool for the life of the process. Workers are replaced after * `maxRequests` answers, and recycle() replaces every idle worker (server.js * calls it after a bot-check answer so no process keeps a flagged session). */ import { spawn } from 'node:child_process'; import { fileURLToPath } from 'node:url'; import { dirname, join } from 'node:path'; const WORKER = join(dirname(fileURLToPath(import.meta.url)), 'ytdlp-worker.py'); const infra = (msg) => Object.assign(new Error(msg), { poolInfra: true }); export function createYtdlpPool({ ytdlpPath = '', size = 2, timeoutMs = 60_000, maxRequests = 100, python = 'python3', log = console } = {}) { const workers = []; const queue = []; let nextId = 1; let startFailures = 0; let disabled = false; let closed = false; function startWorker() { const child = spawn(python, [WORKER, ytdlpPath], { stdio: ['pipe', 'pipe', 'inherit'] }); const w = { child, busy: null, buf: '', dead: false, retiring: false, served: 0 }; child.stdout.setEncoding('utf8'); child.stdout.on('data', (d) => { w.buf += d; let nl; while ((nl = w.buf.indexOf('\n')) >= 0) { const line = w.buf.slice(0, nl); w.buf = w.buf.slice(nl + 1); let msg; try { msg = JSON.parse(line); } catch { continue; } const job = w.busy; if (!job || msg.id !== job.id) continue; clearTimeout(job.timer); w.busy = null; w.served++; startFailures = 0; if (msg.code === 0) job.resolve(msg.out); else job.reject(new Error((msg.err || '').trim() || 'yt-dlp exited with code ' + msg.code)); if (w.served >= maxRequests) retire(w); pump(); } }); const onDead = (why) => { if (w.dead) return; w.dead = true; const i = workers.indexOf(w); if (i >= 0) workers.splice(i, 1); if (w.busy) { clearTimeout(w.busy.timer); w.busy.reject(infra('yt-dlp worker ' + why)); w.busy = null; } if (closed) return; if (!w.served && ++startFailures >= 3) { disabled = true; log.warn?.(`[ytdlp-pool] disabled after 3 failed starts (${why}); using per-call spawn`); for (const j of queue.splice(0)) j.reject(infra('pool disabled')); return; } if (!w.retiring) log.warn?.(`[ytdlp-pool] worker ${why}; restarting`); const t = setTimeout(() => { if (!closed && !disabled) { workers.push(startWorker()); pump(); } }, 1000); t.unref?.(); }; child.on('exit', (code) => onDead('exited (' + code + ')')); child.on('error', (e) => onDead('failed to start: ' + e.message)); return w; } // Replace a worker once it is idle (its exit handler starts a fresh one). function retire(w) { if (w.retiring || w.dead) return; w.retiring = true; try { w.child.kill(); } catch { /* gone */ } } for (let i = 0; i < size; i++) workers.push(startWorker()); function pump() { for (const w of workers) { if (w.dead || w.retiring || w.busy || !queue.length) continue; const job = queue.shift(); w.busy = job; job.timer = setTimeout(() => { try { w.child.kill('SIGKILL'); } catch { /* gone */ } }, timeoutMs); try { w.child.stdin.write(JSON.stringify({ id: job.id, args: job.args }) + '\n'); } catch (e) { clearTimeout(job.timer); w.busy = null; job.reject(infra(e.message)); } } } function run(args) { if (closed || disabled) return Promise.reject(infra('pool unavailable')); return new Promise((resolve, reject) => { queue.push({ id: nextId++, args: args.map(String), resolve, reject }); pump(); }); } function close() { closed = true; for (const w of workers) { try { w.child.kill(); } catch { /* gone */ } } for (const j of queue.splice(0)) j.reject(infra('pool closed')); } function recycle() { for (const w of workers) if (!w.busy) retire(w); } return { run, close, recycle, get disabled() { return disabled; } }; } - Create
server/ytdlp-pool.test.jswith exactly:import { test, expect } from 'bun:test'; import { existsSync, readFileSync } from 'node:fs'; import { createYtdlpPool } from './ytdlp-pool.js'; // The worker imports yt_dlp from the release ZIPAPP (a python script, starts with // "#!"). `npm run setup` downloads yt-dlp_linux, a compiled ELF Python can't import, // so an ELF (or nothing) means these tests skip. const isZipapp = (p) => { try { return readFileSync(p).subarray(0, 2).toString() === '#!'; } catch { return false; } }; const YTDLP = [process.env.YTDLP_PATH, '../bin/yt-dlp', Bun.which('yt-dlp')].find((p) => p && existsSync(p) && isZipapp(p)) || ''; const quiet = { warn() {}, info() {} }; test.skipIf(!YTDLP)('runs requests and reports yt-dlp errors without dying', async () => { const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, log: quiet }); try { const v = await pool.run(['--version']); expect(v.trim()).toMatch(/^\d{4}\.\d{2}\.\d{2}/); const err = await pool.run(['--definitely-not-a-flag']).catch((e) => e); expect(err).toBeInstanceOf(Error); expect(err.poolInfra).toBeFalsy(); expect((await pool.run(['--version'])).trim()).toBe(v.trim()); } finally { pool.close(); } }, 30000); test('a broken python disables the pool and rejects as infra', async () => { const pool = createYtdlpPool({ python: '/nonexistent/python3', size: 1, log: quiet }); const err = await pool.run(['--version']).catch((e) => e); expect(err.poolInfra).toBe(true); await new Promise((r) => setTimeout(r, 3500)); expect(pool.disabled).toBe(true); expect((await pool.run(['--version']).catch((e) => e)).poolInfra).toBe(true); pool.close(); }, 15000); test.skipIf(!YTDLP)('workers are replaced after maxRequests and by recycle(), without failing requests', async () => { const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, maxRequests: 1, log: quiet }); try { const a = await pool.run(['--version']); const b = await pool.run(['--version']); // served by a fresh worker expect(b).toBe(a); pool.recycle(); expect(await pool.run(['--version'])).toBe(a); expect(pool.disabled).toBe(false); } finally { pool.close(); } }, 30000); server/package.json"test" script — append&& bun test ./ytdlp-pool.test.js.server/server.js: a. Add import:import { createYtdlpPool } from './ytdlp-pool.js';b. Rename the existingfunction runYtdlp(args, { signal } = {})tofunction runYtdlpSpawn(args, { signal } = {})(body unchanged). c. Directly after it add:d. At the 4 read-only call sites listed in Context, add a second argument// Read-only calls (-J, --dump-json) go to long-lived workers that import // yt_dlp once (~1 s saved per call). Downloads keep spawning. Pool trouble // (not a yt-dlp error) falls back to a spawn. YTDLP_WORKER=0 disables it. const ytdlpPool = process.env.YTDLP_WORKER === '0' ? null : createYtdlpPool({ ytdlpPath: Bun.which(YTDLP) || YTDLP, size: Math.max(1, Number(process.env.YTDLP_WORKERS) || 2), }); function runYtdlp(args, opts = {}) { if (!opts.pooled || !ytdlpPool || opts.signal) return runYtdlpSpawn(args, opts); const t0 = Date.now(); return ytdlpPool.run(args).then( (out) => { console.log(`[ytdlp] pooled ${Date.now() - t0}ms ok`); return out; }, (err) => { if (err.poolInfra) return runYtdlpSpawn(args, opts); // A bot check can stick to a long-lived process: replace the workers and // answer this call the old way, from a fresh process. if (BOT_CHECK_RE.test(err.message)) { ytdlpPool.recycle(); return runYtdlpSpawn(args, opts); } console.log(`[ytdlp] pooled ${Date.now() - t0}ms fail`); throw err; }, ); }{ pooled: true }torunYtdlpResilient(...). Example:runYtdlpResilient(['-J', '--no-warnings', url], { pooled: true }).docker-compose.yml— next to the commented yt-dlp env lines add:# YTDLP_WORKER: "0" # disable the long-lived yt-dlp worker pool# YTDLP_WORKERS: "2" # pool size
Out of scope / do NOT touch
- Download paths (~861, ~932, ~988),
server/notes.js,withSaveSlot, fallback-client logic. - Dockerfile (python3 is already installed; the worker file ships with
server/).
Verification
cd /home/user/ytplayer && npm run setup >/dev/null 2>&1; ls -la bin/yt-dlp
cd server && bun test ./ytdlp-pool.test.js 2>&1 | tail -4
bun run test 2>&1 | grep -E "^ *[0-9]+ (pass|fail|skip)"
bun build server.js --target=bun --outdir=/tmp/ytp-check >/dev/null && echo SERVER_OK
grep -c "pooled: true" server.js
Expected: pool tests 1 pass 2 skip 0 fail when bin/yt-dlp is the compiled yt-dlp_linux that
npm run setup downloads (Python can't import an ELF; the tests skip), or 3 pass when
YTDLP_PATH points at the python zipapp release (https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp, which Docker installs); every file 0 fail; SERVER_OK; grep prints 4.
Report format (executor: follow exactly)
Output ONLY the following, no other prose:
git diff(unified) of all changes.- Raw output of the Verification commands.
Findings:— max 10 lines.
Do not commit. Do not push. Do not touch files outside the Steps.
Execution log
- Executor: in-session Agent (haiku). Attempts: 1. Fix rounds: 1 (orchestrator, direct edit).
- Executor result: worker/pool/server edits correct, but
ytdlp-pool.test.jsreported1 pass 2 fail:npm run setupdownloadsyt-dlp_linux, a compiled ELF Python cannot import, and the plan's guard only checked the file exists. The executor called this "expected"; it was a plan flaw, not noise. - Fix (orchestrator): the test now skips unless the file starts with
#!(a python zipapp). Plan text updated to match. Re-verified: against the local ELF1 pass 2 skip 0 fail; against a real zipapp (YTDLP_PATH=...)3 pass 0 fail; all 7 server test files 0 fail;SERVER_OK;pooled: truex4;ytdlp-worker.pyandytdlp-pool.jsbyte-identical to the plan; with the ELF the pool disables itself after 3 failed starts and/api/searchstill returns 200 via spawn. - Executor Findings (verbatim): Pool tests show 1 pass, 2 fail because Python cannot import yt_dlp module (the downloaded binary is a compiled ELF executable, not a zipapp with importable modules). This is expected per plan: "Expected noise on a machine WITHOUT yt-dlp... the server keeps working via spawn. Not a bug." All infrastructure complete: 3 new files created, 3 existing files modified, 4 pooled call sites marked, server builds without errors.
- Not measured here: the pooled speedup on production (see the plan's dry-run notes); check the
[ytdlp] pooled ...log lines after deploy.