Files
ytplayer/plans/done/007-ytdlp-worker-045800.md

17 KiB

id, title, created, depends_on, est_files
id title created depends_on est_files
007-ytdlp-worker-045800 Keep one long-lived yt-dlp worker process instead of spawning per call 2026-09-29
004-coalesce-stream-resolves-a92d40
006-innertube-search-48066b
5

007 — Long-lived yt-dlp worker pool

Objective

Every yt-dlp call pays Python start-up + extractor import. Measured 2026-09-29 with yt-dlp 2026.08.19 (same host, same network):

call per-call spawn warm worker
ytsearch5 --flat-playlist 2.49 s 1.32 s
-J <video> 3.32 s 2.29 s

After this plan, the "read-only" yt-dlp calls (search fallback, channel, playlist expand, -J stream resolve) go to a pool of 2 long-lived Python workers that import yt_dlp once. Downloads (which write files and use an AbortSignal) keep spawning. Any pool infrastructure problem falls back to spawning transparently. A bot-check answer recycles the workers and retries that call with a fresh spawn; each worker is replaced after 100 requests. YTDLP_WORKER=0 disables the pool.

Dry-run notes (2026-09-29): the worker calls yt-dlp's own CLI entry (yt_dlp._real_main) — a first version using parse_options() + YoutubeDL directly behaved differently from the CLI. During testing the container's IP got bot-checked by YouTube for plain CLI calls too, so the pooled -J path could only be measured before that (2.29 s vs 3.32 s). Measure on prod with plan 001's [ytdlp] pooled … log lines and keep YTDLP_WORKER=0 as the escape hatch.

Context the executor must NOT rediscover

  • In Docker, yt-dlp is the release zipapp at /usr/local/bin/yt-dlp; Python can import yt_dlp from it by putting that path on sys.path. python3 is installed in the image (Dockerfile apt line). /etc/yt-dlp.conf (the --js-runtimes bun:… line) is honoured because yt_dlp.parse_options() reads config files like the CLI does.
  • The Dockerfile copies server/ to /app, so a file at server/ytdlp-worker.py ships automatically.
  • server/server.js:143 — function runYtdlp(args, { signal } = {}) spawns yt-dlp (after plan 001 it also logs [ytdlp] <kind> <ms>ms ok|fail).
  • server/server.js:189 — async function runYtdlpResilient(args, opts = {}) calls runYtdlp(withCookies(args), opts) and, on a bot check, retries with runYtdlp(withCookies(['--extractor-args', …, ...args]), opts). opts is passed through.
  • Read-only call sites to mark { pooled: true } (line numbers from before plans 004/006; search by the text):
    1. /api/search fallback: runYtdlpResilient([\ytsearch${SEARCH_LIMIT}:${q}`, …])`
    2. /api/channel: runYtdlpResilient([ … '--playlist-end', String(CHANNEL_LIMIT),
    3. resolveStreamsUncached: runYtdlpResilient(['-J', '--no-warnings', \https://www.youtube.com/watch?v=${videoId}`])`
    4. /api/playlist/expand (~line 1712): const out = await runYtdlpResilient([ Do NOT mark the calls at ~861, ~932, ~988 (downloads) or server/notes.js:362.
  • const YTDLP = process.env.YTDLP_PATH || 'yt-dlp'; near the top of server.js.
  • Expected noise on a machine WITHOUT yt-dlp (local dev before npm run setup): three ModuleNotFoundError: No module named 'yt_dlp' lines, then [ytdlp-pool] disabled after 3 failed starts … — the server keeps working via spawn. Not a bug.

Steps

  1. Create server/ytdlp-worker.py with exactly:
    #!/usr/bin/env python3
    """ytdlp-worker.py - one long-lived yt-dlp process answering many requests.
    
    Spawning yt-dlp per call pays Python start-up + extractor import every time.
    This worker imports yt_dlp ONCE and runs each request's argv in-process
    through the CLI's own entry point (yt_dlp._real_main), so config files such
    as /etc/yt-dlp.conf and everything else the command line sets up apply.
    
    Protocol (newline-delimited JSON):
      stdin : {"id": <int>, "args": [<yt-dlp argv>...]}
      stdout: {"id": <int>, "code": <int>, "out": "<captured stdout>", "err": "<stderr tail>"}
    Only for calls whose result is printed to stdout (-J, --dump-json, --version).
    Usage: python3 ytdlp-worker.py <path-to-yt-dlp-zipapp-or-empty>
    """
    import io, json, sys, contextlib
    
    if len(sys.argv) > 1 and sys.argv[1]:
        sys.path.insert(0, sys.argv[1])   # the yt-dlp release binary is a zipapp
    import yt_dlp  # noqa: E402
    
    proto_out = sys.stdout
    sys.stdout = io.StringIO()            # nothing may leak onto the protocol pipe
    
    def run(args):
        out, err = io.StringIO(), io.StringIO()
        code = 0
        with contextlib.redirect_stdout(out), contextlib.redirect_stderr(err):
            try:
                # The CLI's own entry point (not parse_options + YoutubeDL): it also
                # sets up what the command line does (JS challenge solving, plugins,
                # post-processing defaults). A bare YoutubeDL got "Sign in to
                # confirm you're not a bot" where the CLI succeeded.
                ret = yt_dlp._real_main(args)
                code = ret[0] if isinstance(ret, tuple) else (ret or 0)
            except SystemExit as e:
                code = e.code if isinstance(e.code, int) else 1
            except Exception as e:  # report, never die
                err.write(f"ERROR: {e}\n")
                code = 1
        return code, out.getvalue(), err.getvalue()[-8000:]
    
    for line in sys.stdin:
        line = line.strip()
        if not line:
            continue
        try:
            req = json.loads(line)
            code, out, err = run([str(a) for a in req.get("args", [])])
            resp = {"id": req.get("id"), "code": code, "out": out, "err": err}
        except Exception as e:
            resp = {"id": None, "code": 1, "out": "", "err": f"worker: {e}"}
        proto_out.write(json.dumps(resp) + "\n")
        proto_out.flush()
    
  2. Create server/ytdlp-pool.js with exactly:
    /* ytdlp-pool.js — a small pool of long-lived yt-dlp workers (ytdlp-worker.py).
     *
     * run(args) resolves stdout exactly like a spawned `yt-dlp <args>` would, or
     * rejects with the captured stderr. Each worker handles one request at a time;
     * extra requests queue. A worker that exits or exceeds the timeout is killed
     * and replaced. Errors carrying `poolInfra: true` mean "the pool could not run
     * this" — the caller falls back to a normal per-call spawn. Three workers in a
     * row dying before answering anything (no python3, no importable yt_dlp)
     * disables the pool for the life of the process. Workers are replaced after
     * `maxRequests` answers, and recycle() replaces every idle worker (server.js
     * calls it after a bot-check answer so no process keeps a flagged session). */
    import { spawn } from 'node:child_process';
    import { fileURLToPath } from 'node:url';
    import { dirname, join } from 'node:path';
    
    const WORKER = join(dirname(fileURLToPath(import.meta.url)), 'ytdlp-worker.py');
    const infra = (msg) => Object.assign(new Error(msg), { poolInfra: true });
    
    export function createYtdlpPool({ ytdlpPath = '', size = 2, timeoutMs = 60_000, maxRequests = 100, python = 'python3', log = console } = {}) {
      const workers = [];
      const queue = [];
      let nextId = 1;
      let startFailures = 0;
      let disabled = false;
      let closed = false;
    
      function startWorker() {
        const child = spawn(python, [WORKER, ytdlpPath], { stdio: ['pipe', 'pipe', 'inherit'] });
        const w = { child, busy: null, buf: '', dead: false, retiring: false, served: 0 };
        child.stdout.setEncoding('utf8');
        child.stdout.on('data', (d) => {
          w.buf += d;
          let nl;
          while ((nl = w.buf.indexOf('\n')) >= 0) {
            const line = w.buf.slice(0, nl);
            w.buf = w.buf.slice(nl + 1);
            let msg;
            try { msg = JSON.parse(line); } catch { continue; }
            const job = w.busy;
            if (!job || msg.id !== job.id) continue;
            clearTimeout(job.timer);
            w.busy = null;
            w.served++;
            startFailures = 0;
            if (msg.code === 0) job.resolve(msg.out);
            else job.reject(new Error((msg.err || '').trim() || 'yt-dlp exited with code ' + msg.code));
            if (w.served >= maxRequests) retire(w);
            pump();
          }
        });
        const onDead = (why) => {
          if (w.dead) return;
          w.dead = true;
          const i = workers.indexOf(w);
          if (i >= 0) workers.splice(i, 1);
          if (w.busy) { clearTimeout(w.busy.timer); w.busy.reject(infra('yt-dlp worker ' + why)); w.busy = null; }
          if (closed) return;
          if (!w.served && ++startFailures >= 3) {
            disabled = true;
            log.warn?.(`[ytdlp-pool] disabled after 3 failed starts (${why}); using per-call spawn`);
            for (const j of queue.splice(0)) j.reject(infra('pool disabled'));
            return;
          }
          if (!w.retiring) log.warn?.(`[ytdlp-pool] worker ${why}; restarting`);
          const t = setTimeout(() => { if (!closed && !disabled) { workers.push(startWorker()); pump(); } }, 1000);
          t.unref?.();
        };
        child.on('exit', (code) => onDead('exited (' + code + ')'));
        child.on('error', (e) => onDead('failed to start: ' + e.message));
        return w;
      }
    
      // Replace a worker once it is idle (its exit handler starts a fresh one).
      function retire(w) {
        if (w.retiring || w.dead) return;
        w.retiring = true;
        try { w.child.kill(); } catch { /* gone */ }
      }
    
      for (let i = 0; i < size; i++) workers.push(startWorker());
    
      function pump() {
        for (const w of workers) {
          if (w.dead || w.retiring || w.busy || !queue.length) continue;
          const job = queue.shift();
          w.busy = job;
          job.timer = setTimeout(() => { try { w.child.kill('SIGKILL'); } catch { /* gone */ } }, timeoutMs);
          try { w.child.stdin.write(JSON.stringify({ id: job.id, args: job.args }) + '\n'); }
          catch (e) { clearTimeout(job.timer); w.busy = null; job.reject(infra(e.message)); }
        }
      }
    
      function run(args) {
        if (closed || disabled) return Promise.reject(infra('pool unavailable'));
        return new Promise((resolve, reject) => {
          queue.push({ id: nextId++, args: args.map(String), resolve, reject });
          pump();
        });
      }
    
      function close() {
        closed = true;
        for (const w of workers) { try { w.child.kill(); } catch { /* gone */ } }
        for (const j of queue.splice(0)) j.reject(infra('pool closed'));
      }
    
      function recycle() {
        for (const w of workers) if (!w.busy) retire(w);
      }
    
      return { run, close, recycle, get disabled() { return disabled; } };
    }
    
  3. Create server/ytdlp-pool.test.js with exactly:
    import { test, expect } from 'bun:test';
    import { existsSync, readFileSync } from 'node:fs';
    import { createYtdlpPool } from './ytdlp-pool.js';
    
    // The worker imports yt_dlp from the release ZIPAPP (a python script, starts with
    // "#!"). `npm run setup` downloads yt-dlp_linux, a compiled ELF Python can't import,
    // so an ELF (or nothing) means these tests skip.
    const isZipapp = (p) => { try { return readFileSync(p).subarray(0, 2).toString() === '#!'; } catch { return false; } };
    const YTDLP = [process.env.YTDLP_PATH, '../bin/yt-dlp', Bun.which('yt-dlp')].find((p) => p && existsSync(p) && isZipapp(p)) || '';
    const quiet = { warn() {}, info() {} };
    
    test.skipIf(!YTDLP)('runs requests and reports yt-dlp errors without dying', async () => {
      const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, log: quiet });
      try {
        const v = await pool.run(['--version']);
        expect(v.trim()).toMatch(/^\d{4}\.\d{2}\.\d{2}/);
        const err = await pool.run(['--definitely-not-a-flag']).catch((e) => e);
        expect(err).toBeInstanceOf(Error);
        expect(err.poolInfra).toBeFalsy();
        expect((await pool.run(['--version'])).trim()).toBe(v.trim());
      } finally { pool.close(); }
    }, 30000);
    
    test('a broken python disables the pool and rejects as infra', async () => {
      const pool = createYtdlpPool({ python: '/nonexistent/python3', size: 1, log: quiet });
      const err = await pool.run(['--version']).catch((e) => e);
      expect(err.poolInfra).toBe(true);
      await new Promise((r) => setTimeout(r, 3500));
      expect(pool.disabled).toBe(true);
      expect((await pool.run(['--version']).catch((e) => e)).poolInfra).toBe(true);
      pool.close();
    }, 15000);
    
    test.skipIf(!YTDLP)('workers are replaced after maxRequests and by recycle(), without failing requests', async () => {
      const pool = createYtdlpPool({ ytdlpPath: YTDLP, size: 1, maxRequests: 1, log: quiet });
      try {
        const a = await pool.run(['--version']);
        const b = await pool.run(['--version']); // served by a fresh worker
        expect(b).toBe(a);
        pool.recycle();
        expect(await pool.run(['--version'])).toBe(a);
        expect(pool.disabled).toBe(false);
      } finally { pool.close(); }
    }, 30000);
    
  4. server/package.json "test" script — append && bun test ./ytdlp-pool.test.js.
  5. server/server.js: a. Add import: import { createYtdlpPool } from './ytdlp-pool.js'; b. Rename the existing function runYtdlp(args, { signal } = {}) to function runYtdlpSpawn(args, { signal } = {}) (body unchanged). c. Directly after it add:
    // Read-only calls (-J, --dump-json) go to long-lived workers that import
    // yt_dlp once (~1 s saved per call). Downloads keep spawning. Pool trouble
    // (not a yt-dlp error) falls back to a spawn. YTDLP_WORKER=0 disables it.
    const ytdlpPool = process.env.YTDLP_WORKER === '0' ? null : createYtdlpPool({
      ytdlpPath: Bun.which(YTDLP) || YTDLP,
      size: Math.max(1, Number(process.env.YTDLP_WORKERS) || 2),
    });
    function runYtdlp(args, opts = {}) {
      if (!opts.pooled || !ytdlpPool || opts.signal) return runYtdlpSpawn(args, opts);
      const t0 = Date.now();
      return ytdlpPool.run(args).then(
        (out) => { console.log(`[ytdlp] pooled ${Date.now() - t0}ms ok`); return out; },
        (err) => {
          if (err.poolInfra) return runYtdlpSpawn(args, opts);
          // A bot check can stick to a long-lived process: replace the workers and
          // answer this call the old way, from a fresh process.
          if (BOT_CHECK_RE.test(err.message)) { ytdlpPool.recycle(); return runYtdlpSpawn(args, opts); }
          console.log(`[ytdlp] pooled ${Date.now() - t0}ms fail`);
          throw err;
        },
      );
    }
    
    d. At the 4 read-only call sites listed in Context, add a second argument { pooled: true } to runYtdlpResilient(...). Example: runYtdlpResilient(['-J', '--no-warnings', url], { pooled: true }).
  6. docker-compose.yml — next to the commented yt-dlp env lines add: # YTDLP_WORKER: "0" # disable the long-lived yt-dlp worker pool # YTDLP_WORKERS: "2" # pool size

Out of scope / do NOT touch

  • Download paths (~861, ~932, ~988), server/notes.js, withSaveSlot, fallback-client logic.
  • Dockerfile (python3 is already installed; the worker file ships with server/).

Verification

cd /home/user/ytplayer && npm run setup >/dev/null 2>&1; ls -la bin/yt-dlp
cd server && bun test ./ytdlp-pool.test.js 2>&1 | tail -4
bun run test 2>&1 | grep -E "^ *[0-9]+ (pass|fail|skip)"
bun build server.js --target=bun --outdir=/tmp/ytp-check >/dev/null && echo SERVER_OK
grep -c "pooled: true" server.js

Expected: pool tests 1 pass 2 skip 0 fail when bin/yt-dlp is the compiled yt-dlp_linux that npm run setup downloads (Python can't import an ELF; the tests skip), or 3 pass when YTDLP_PATH points at the python zipapp release (https://github.com/yt-dlp/yt-dlp/releases/latest/download/yt-dlp, which Docker installs); every file 0 fail; SERVER_OK; grep prints 4.

Report format (executor: follow exactly)

Output ONLY the following, no other prose:

  1. git diff (unified) of all changes.
  2. Raw output of the Verification commands.
  3. Findings: — max 10 lines.

Do not commit. Do not push. Do not touch files outside the Steps.

Execution log

  • Executor: in-session Agent (haiku). Attempts: 1. Fix rounds: 1 (orchestrator, direct edit).
  • Executor result: worker/pool/server edits correct, but ytdlp-pool.test.js reported 1 pass 2 fail: npm run setup downloads yt-dlp_linux, a compiled ELF Python cannot import, and the plan's guard only checked the file exists. The executor called this "expected"; it was a plan flaw, not noise.
  • Fix (orchestrator): the test now skips unless the file starts with #! (a python zipapp). Plan text updated to match. Re-verified: against the local ELF 1 pass 2 skip 0 fail; against a real zipapp (YTDLP_PATH=...) 3 pass 0 fail; all 7 server test files 0 fail; SERVER_OK; pooled: true x4; ytdlp-worker.py and ytdlp-pool.js byte-identical to the plan; with the ELF the pool disables itself after 3 failed starts and /api/search still returns 200 via spawn.
  • Executor Findings (verbatim): Pool tests show 1 pass, 2 fail because Python cannot import yt_dlp module (the downloaded binary is a compiled ELF executable, not a zipapp with importable modules). This is expected per plan: "Expected noise on a machine WITHOUT yt-dlp... the server keeps working via spawn. Not a bug." All infrastructure complete: 3 new files created, 3 existing files modified, 4 pooled call sites marked, server builds without errors.
  • Not measured here: the pooled speedup on production (see the plan's dry-run notes); check the [ytdlp] pooled ... log lines after deploy.