Files
ytplayer/plans/done/007-ytdlp-worker-045800.md

351 lines
17 KiB
Markdown

---
id: 007-ytdlp-worker-045800
title: Keep one long-lived yt-dlp worker process instead of spawning per call
created: 2026-09-29
depends_on: [004-coalesce-stream-resolves-a92d40, 006-innertube-search-48066b]
est_files: 5
---
# 007 — Long-lived yt-dlp worker pool
## Objective
Every yt-dlp call pays Python start-up + extractor import. Measured 2026-09-29 with
yt-dlp 2026.08.19 (same host, same network):
| call | per-call spawn | warm worker |
|------|----------------|-------------|
| `ytsearch5 --flat-playlist` | 2.49 s | 1.32 s |
| `-J <video>` | 3.32 s | 2.29 s |
After this plan, the "read-only" yt-dlp calls (search fallback, channel, playlist
expand, `-J` stream resolve) go to a pool of 2 long-lived Python workers that import
`yt_dlp` once. Downloads (which write files and use an AbortSignal) keep spawning.
Any pool infrastructure problem falls back to spawning transparently. A bot-check answer
recycles the workers and retries that call with a fresh spawn; each worker is replaced after
100 requests. `YTDLP_WORKER=0` disables the pool.
Dry-run notes (2026-09-29): the worker calls yt-dlp's own CLI entry (`yt_dlp._real_main`) — a
first version using `parse_options()` + `YoutubeDL` directly behaved differently from the CLI.
During testing the container's IP got bot-checked by YouTube for plain CLI calls too, so the
pooled `-J` path could only be measured before that (2.29 s vs 3.32 s). Measure on prod with
plan 001's `[ytdlp] pooled …` log lines and keep `YTDLP_WORKER=0` as the escape hatch.
## Context the executor must NOT rediscover
- In Docker, yt-dlp is the release **zipapp** at `/usr/local/bin/yt-dlp`; Python can import
`yt_dlp` from it by putting that path on `sys.path`. `python3` is installed in the image
(Dockerfile apt line). `/etc/yt-dlp.conf` (the `--js-runtimes bun:…` line) is honoured
because `yt_dlp.parse_options()` reads config files like the CLI does.
- The Dockerfile copies `server/` to `/app`, so a file at `server/ytdlp-worker.py` ships
automatically.
- `server/server.js:143` — `function runYtdlp(args, { signal } = {})` spawns yt-dlp
(after plan 001 it also logs `[ytdlp] <kind> <ms>ms ok|fail`).
- `server/server.js:189` — `async function runYtdlpResilient(args, opts = {})` calls
`runYtdlp(withCookies(args), opts)` and, on a bot check, retries with
`runYtdlp(withCookies(['--extractor-args', …, ...args]), opts)`. `opts` is passed through.
- Read-only call sites to mark `{ pooled: true }` (line numbers from before plans 004/006;
search by the text):
1. `/api/search` fallback: `runYtdlpResilient([\`ytsearch${SEARCH_LIMIT}:${q}\`, …])`
2. `/api/channel`: `runYtdlpResilient([` … `'--playlist-end', String(CHANNEL_LIMIT),`
3. `resolveStreamsUncached`: `runYtdlpResilient(['-J', '--no-warnings', \`https://www.youtube.com/watch?v=${videoId}\`])`
4. `/api/playlist/expand` (~line 1712): `const out = await runYtdlpResilient([`
Do NOT mark the calls at ~861, ~932, ~988 (downloads) or `server/notes.js:362`.
- `const YTDLP = process.env.YTDLP_PATH || 'yt-dlp';` near the top of server.js.
- Expected noise on a machine WITHOUT yt-dlp (local dev before `npm run setup`): three
`ModuleNotFoundError: No module named 'yt_dlp'` lines, then
`[ytdlp-pool] disabled after 3 failed starts …` — the server keeps working via spawn. Not a bug.
## Steps
1. Create `server/ytdlp-worker.py` with exactly:
```python
#!/usr/bin/env python3
"""ytdlp-worker.py - one long-lived yt-dlp process answering many requests.
Spawning yt-dlp per call pays Python start-up + extractor import every time.
This worker imports yt_dlp ONCE and runs each request's argv in-process
through the CLI's own entry point (yt_dlp._real_main), so config files such
as /etc/yt-dlp.conf and everything else the command line sets up apply.
Protocol (newline-delimited JSON):
stdin : {"id": <int>, "args": [<yt-dlp argv>...]}
stdout: {"id": <int>, "code": <int>, "out": "<captured stdout>", "err": "<stderr tail>"}
Only for calls whose result is printed to stdout (-J, --dump-json, --version).
Usage: python3 ytdlp-worker.py <path-to-yt-dlp-zipapp-or-empty>
"""
import io, json, sys, contextlib
if len(sys.argv) > 1 and sys.argv[1]:
sys.path.insert(0, sys.argv[1]) # the yt-dlp release binary is a zipapp
import yt_dlp # noqa: E402
proto_out = sys.stdout
sys.stdout = io.StringIO() # nothing may leak onto the protocol pipe
def run(args):
out, err = io.StringIO(), io.StringIO()
code = 0
with contextlib.redirect_stdout(out), contextlib.redirect_stderr(err):
try:
# The CLI's own entry point (not parse_options + YoutubeDL): it also
# sets up what the command line does (JS challenge solving, plugins,
# post-processing defaults). A bare YoutubeDL got "Sign in to
# confirm you're not a bot" where the CLI succeeded.
ret = yt_dlp._real_main(args)
code = ret[0] if isinstance(ret, tuple) else (ret or 0)
except SystemExit as e:
code = e.code if isinstance(e.code, int) else 1
except Exception as e: # report, never die
err.write(f"ERROR: {e}\n")
code = 1
return code, out.getvalue(), err.getvalue()[-8000:]
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
req = json.loads(line)
code, out, err = run([str(a) for a in req.get("args", [])])
resp = {"id": req.get("id"), "code": code, "out": out, "err": err}
except Exception as e:
resp = {"id": None, "code": 1, "out": "", "err": f"worker: {e}"}
proto_out.write(json.dumps(resp) + "\n")
proto_out.flush()
```
2. Create `server/ytdlp-pool.js` with exactly:
```js
/* ytdlp-pool.js — a small pool of long-lived yt-dlp workers (ytdlp-worker.py).
*
* run(args) resolves stdout exactly like a spawned `yt-dlp <args>` would, or
* rejects with the captured stderr. Each worker handles one request at a time;
* extra requests queue. A worker that exits or exceeds the timeout is killed
* and replaced. Errors carrying `poolInfra: true` mean "the pool could not run
* this" — the caller falls back to a normal per-call spawn. Three workers in a
* row dying before answering anything (no python3, no importable yt_dlp)
* disables the pool for the life of the process. Workers are replaced after
* `maxRequests` answers, and recycle() replaces every idle worker (server.js
* calls it after a bot-check answer so no process keeps a flagged session). */
import { spawn } from 'node:child_process';
import { fileURLToPath } from 'node:url';
import { dirname, join } from 'node:path';
const WORKER = join(dirname(fileURLToPath(import.meta.url)), 'ytdlp-worker.py');
const infra = (msg) => Object.assign(new Error(msg), { poolInfra: true });
export function createYtdlpPool({ ytdlpPath = '', size = 2, timeoutMs = 60_000, maxRequests = 100, python = 'python3', log = console } = {}) {
const workers = [];
const queue = [];
let nextId = 1;
let startFailures = 0;
let disabled = false;
let closed = false;
function startWorker() {
const child = spawn(python, [WORKER, ytdlpPath], { stdio: ['pipe', 'pipe', 'inherit'] });
const w = { child, busy: null, buf: '', dead: false, retiring: false, served: 0 };
child.stdout.setEncoding('utf8');
child.stdout.on('data', (d) => {
w.buf += d;
let nl;
while ((nl = w.buf.indexOf('\n')) >= 0) {
const line = w.buf.slice(0, nl);
w.buf = w.buf.slice(nl + 1);
let msg;
try { msg = JSON.parse(line); } catch { continue; }
const job = w.busy;
if (!job || msg.id !== job.id) continue;
clearTimeout(job.timer);
w.busy = null;
w.served++;
startFailures = 0;
if (msg.code === 0) job.resolve(msg.out);
else job.reject(new Error((msg.err || '').trim() || 'yt-dlp exited with code ' + msg.code));
if (w.served >= maxRequests) retire(w);
pump();
}
});
const onDead = (why) => {
if (w.dead) return;
w.dead = true;
const i = workers.indexOf(w);
if (i >= 0) workers.splice(i, 1);
if (w.busy) { clearTimeout(w.busy.timer); w.busy.reject(infra('yt-dlp worker ' + why)); w.busy = null; }
if (closed) return;
if (!w.served && ++startFailures >= 3) {
disabled = true;
log.warn?.(`[ytdlp-pool] disabled after 3 failed starts (${why}); using per-call spawn`);
for (const j of queue.splice(0)) j.reject(infra('pool disabled'));
return;
}
if (!w.retiring) log.warn?.(`[ytdlp-pool] worker ${why}; restarting`);
const t = setTimeout(() => { if (!closed && !disabled) { workers.push(startWorker()); pump(); } }, 1000);
t.unref?.();
};
child.on('exit', (code) => onDead('exited (' + code + ')'));
child.on('error', (e) => onDead('failed to start: ' + e.message));
return w;
}
// Replace a worker once it is idle (its exit handler starts a fresh one).
function retire(w) {
if (w.retiring || w.dead) return;
w.retiring = true;
try { w.child.kill(); } catch { /* gone */ }
}
for (let i = 0; i < size; i++) workers.push(startWorker());
function pump() {
for (const w of workers) {
if (w.dead || w.retiring || w.busy || !queue.length) continue;
const job = queue.shift();
w.busy = job;
job.timer = setTimeout(() => { try { w.child.kill('SIGKILL'); } catch { /* gone */ } }, timeoutMs);
try { w.child.stdin.write(JSON.stringify({ id: job.id, args: job.args }) + '\n'); }
catch (e) { clearTimeout(job.timer); w.busy = null; job.reject(infra(e.message)); }
}
}
function run(args) {
if (closed || disabled) return Promise.reject(infra('pool unavailable'));
return new Promise((resolve, reject) => {
queue.push({ id: nextId++, args: args.map(String), resolve, reject });
pump();
});
}
function close() {
closed = true;
for (const w of workers) { try { w.child.kill(); } catch { /* gone */ } }
for (const j of queue.splice(0)) j.reject(infra('pool closed'));
}
function recycle() {
for (const w of workers) if (!w.busy) retire(w);
}
return { run, close, recycle, get disabled() { return disabled; } };
}
```
3. Create `server/ytdlp-pool.test.js` with exactly:
```js
import { test, expect } from 'bun:test';
import { existsSync, readFileSync } from 'node:fs';
import { createYtdlpPool } from './ytdlp-pool.js';
// The worker imports yt_dlp from the release ZIPAPP (a python script, starts with
// "#!"). `npm run setup` downloads yt-dlp_linux, a compiled ELF Python can't import,
// so an ELF (or nothing) means these tests skip.
const isZipapp = (p) => { try { return readFileSync(p).subarray(0, 2).toString() === '#!'; } catch { return false; } };
const YTDLP = [process.env.YTDLP_PATH, '../bin/yt-dlp', Bun.which('yt-dlp')].find((p) => p && existsSync(p) && isZipapp(p)) || '';
const quiet = { warn() {}, info() {} };
test.skipIf(!YTDLP)('runs requests and reports yt-dlp errors without dying', async () => {
const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, log: quiet });
try {
const v = await pool.run(['--version']);
expect(v.trim()).toMatch(/^\d{4}\.\d{2}\.\d{2}/);
const err = await pool.run(['--definitely-not-a-flag']).catch((e) => e);
expect(err).toBeInstanceOf(Error);
expect(err.poolInfra).toBeFalsy();
expect((await pool.run(['--version'])).trim()).toBe(v.trim());
} finally { pool.close(); }
}, 30000);
test('a broken python disables the pool and rejects as infra', async () => {
const pool = createYtdlpPool({ python: '/nonexistent/python3', size: 1, log: quiet });
const err = await pool.run(['--version']).catch((e) => e);
expect(err.poolInfra).toBe(true);
await new Promise((r) => setTimeout(r, 3500));
expect(pool.disabled).toBe(true);
expect((await pool.run(['--version']).catch((e) => e)).poolInfra).toBe(true);
pool.close();
}, 15000);
test.skipIf(!YTDLP)('workers are replaced after maxRequests and by recycle(), without failing requests', async () => {
const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, maxRequests: 1, log: quiet });
try {
const a = await pool.run(['--version']);
const b = await pool.run(['--version']); // served by a fresh worker
expect(b).toBe(a);
pool.recycle();
expect(await pool.run(['--version'])).toBe(a);
expect(pool.disabled).toBe(false);
} finally { pool.close(); }
}, 30000);
```
4. `server/package.json` "test" script — append ` && bun test ./ytdlp-pool.test.js`.
5. `server/server.js`:
a. Add import: `import { createYtdlpPool } from './ytdlp-pool.js';`
b. Rename the existing `function runYtdlp(args, { signal } = {})` to
`function runYtdlpSpawn(args, { signal } = {})` (body unchanged).
c. Directly after it add:
```js
// Read-only calls (-J, --dump-json) go to long-lived workers that import
// yt_dlp once (~1 s saved per call). Downloads keep spawning. Pool trouble
// (not a yt-dlp error) falls back to a spawn. YTDLP_WORKER=0 disables it.
const ytdlpPool = process.env.YTDLP_WORKER === '0' ? null : createYtdlpPool({
ytdlpPath: Bun.which(YTDLP) || YTDLP,
size: Math.max(1, Number(process.env.YTDLP_WORKERS) || 2),
});
function runYtdlp(args, opts = {}) {
if (!opts.pooled || !ytdlpPool || opts.signal) return runYtdlpSpawn(args, opts);
const t0 = Date.now();
return ytdlpPool.run(args).then(
(out) => { console.log(`[ytdlp] pooled ${Date.now() - t0}ms ok`); return out; },
(err) => {
if (err.poolInfra) return runYtdlpSpawn(args, opts);
// A bot check can stick to a long-lived process: replace the workers and
// answer this call the old way, from a fresh process.
if (BOT_CHECK_RE.test(err.message)) { ytdlpPool.recycle(); return runYtdlpSpawn(args, opts); }
console.log(`[ytdlp] pooled ${Date.now() - t0}ms fail`);
throw err;
},
);
}
```
d. At the 4 read-only call sites listed in Context, add a second argument `{ pooled: true }`
to `runYtdlpResilient(...)`. Example: `runYtdlpResilient(['-J', '--no-warnings', url], { pooled: true })`.
6. `docker-compose.yml` — next to the commented yt-dlp env lines add:
`# YTDLP_WORKER: "0" # disable the long-lived yt-dlp worker pool`
`# YTDLP_WORKERS: "2" # pool size`
## Out of scope / do NOT touch
- Download paths (~861, ~932, ~988), `server/notes.js`, `withSaveSlot`, fallback-client logic.
- Dockerfile (python3 is already installed; the worker file ships with `server/`).
## Verification
```bash
cd /home/user/ytplayer && npm run setup >/dev/null 2>&1; ls -la bin/yt-dlp
cd server && bun test ./ytdlp-pool.test.js 2>&1 | tail -4
bun run test 2>&1 | grep -E "^ *[0-9]+ (pass|fail|skip)"
bun build server.js --target=bun --outdir=/tmp/ytp-check >/dev/null && echo SERVER_OK
grep -c "pooled: true" server.js
```
Expected: pool tests `1 pass 2 skip 0 fail` when `bin/yt-dlp` is the compiled `yt-dlp_linux` that
`npm run setup` downloads (Python can't import an ELF; the tests skip), or `3 pass` when
`YTDLP_PATH` points at the python zipapp release (`https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp`, which Docker installs); every file `0 fail`; `SERVER_OK`; grep prints `4`.
## Report format (executor: follow exactly)
Output ONLY the following, no other prose:
1. `git diff` (unified) of all changes.
2. Raw output of the Verification commands.
3. `Findings:` — max 10 lines.
Do not commit. Do not push. Do not touch files outside the Steps.
## Execution log
- Executor: in-session Agent (haiku). Attempts: 1. Fix rounds: 1 (orchestrator, direct edit).
- Executor result: worker/pool/server edits correct, but `ytdlp-pool.test.js` reported `1 pass 2 fail`: `npm run setup` downloads `yt-dlp_linux`, a compiled ELF Python cannot import, and the plan's guard only checked the file exists. The executor called this "expected"; it was a plan flaw, not noise.
- Fix (orchestrator): the test now skips unless the file starts with `#!` (a python zipapp). Plan text updated to match. Re-verified: against the local ELF `1 pass 2 skip 0 fail`; against a real zipapp (`YTDLP_PATH=...`) `3 pass 0 fail`; all 7 server test files 0 fail; `SERVER_OK`; `pooled: true` x4; `ytdlp-worker.py` and `ytdlp-pool.js` byte-identical to the plan; with the ELF the pool disables itself after 3 failed starts and `/api/search` still returns 200 via spawn.
- Executor Findings (verbatim): Pool tests show 1 pass, 2 fail because Python cannot import yt_dlp module (the downloaded binary is a compiled ELF executable, not a zipapp with importable modules). This is expected per plan: "Expected noise on a machine WITHOUT yt-dlp... the server keeps working via spawn. Not a bug." All infrastructure complete: 3 new files created, 3 existing files modified, 4 pooled call sites marked, server builds without errors.
- Not measured here: the pooled speedup on production (see the plan's dry-run notes); check the `[ytdlp] pooled ...` log lines after deploy.