# Piano transcription Commit 258aad8 added ideas, not runtime code. The implementation is now opt-in under Settings → Piano transcription. ## What runs where - **Browser:** `frontend/piano-engine.mjs` imports pinned Spotify Basic Pitch 1.0.1 and TensorFlow.js 4.22.0 only after “Transcribe in browser”. It resamples to mono 22,050 Hz in OfflineAudioContext, prefers WebGPU, and falls back to WebGL/CPU. Dependencies load from esm.sh; model/weights load from jsDelivr. Neither the engine nor model is in the service-worker install shell. No audio is uploaded by this path. Processing still needs a network on first use. File/device/server-cached audio is limited to 100 MB and 15 minutes. Basic Pitch is instrument-agnostic, not piano stem separation; mixtures may produce inaccurate notes. - **Device cache:** IndexedDB `ytp-piano/notes`, per video id, stores the engine label and compact JSON sequence. Export produces `.notes.json`. Each note is `{pitch,start,end,velocity,hand}`. Timing is in seconds; velocity is 1–127. Missing hands use a middle-C split (pitch < 60 = left), which is a visual aid, not reliable fingering. Settings shows an 88-key synced visualizer and a manual practice queue of 8-second A–B loops. - **Server:** optional `server/piano.js` queues authenticated jobs in SQLite; `scripts/piano/worker.py` downloads server media on the private compose network and runs ByteDance's `piano_transcription_inference` on CPU. This does not run in the web server. Model loading happens at the first claimed job. The worker is off by default and the compose `piano` profile is inactive by default. The checkpoint is about 165 MB; runtime memory needs substantially more (container limit 6 GB). ## Enable the optional worker Configure a dedicated random `PIANO_WORKER_TOKEN` of at least 24 characters and `PIANO_WORKER_ENABLED=1` for the app and worker. Enable the compose `piano` profile and build/start the worker during a separately approved deployment. It shares the existing private `lyrics` network, keeps its checkpoint in `piano-models`, and is limited to two CPUs. Nothing is pushed or deployed by these changes. Use an admin API token in the Settings server-request field (kept only in the mounted UI, not persisted), or call the contract directly. Save the song on the server first. “Load server result” retrieves the finished sequence and stores it on the device. Worker jobs survive server restarts, have heartbeat leases and three attempts; expired workers cannot publish. Queue capacity is ten. Replacing a server media generation marks its old sequence stale. ## API contract - `GET /api/media/:id/piano`: public `{ok,enabled,job}`; `job` is null or `{videoId,status,notes,error,updatedAt}`. Notes are available only after completion. Status: queued/running/ready/failed/stale. - `POST /api/media/:id/piano`: existing admin session/API-token authentication. Returns an existing valid job or queues one (202). Worker off → 503; uncached source → 409; duration over 15 minutes → 422; queue full → 429. - `POST /api/piano-worker/claim`: dedicated Bearer worker token; returns a job with lease and server-relative audioPath, or null. Claims are atomic. - `POST /api/piano-worker/:id`: dedicated worker token plus JSON `{lease,action}`. Actions: heartbeat; complete with `notes`; fail with `error`. Invalid/expired leases → 409. Results are bounded and validated before saving. The model path is based on [Spotify's browser API](https://github.com/spotify/basic-pitch-ts) and [ByteDance's inference package](https://github.com/qiuqiangkong/piano_transcription_inference). Their upstream licences apply. The browser's TFJS dependency override is pinned and tested with a synthetic WAV; review GPU performance and accuracy with real recordings. ## Review on devices Verify no model request before Transcribe, GPU/CPU fallback, mobile memory limits, cache survival after reload, JSON export, keyboard alignment with seeking, and A–B practice loop timing. Run the optional worker against a short known piano recording before enabling it for users; heavy-model/container inference is not covered by ordinary unit tests. This version does not provide model-based stem separation, score notation, pedal transcription display, or MIDI wait-for-note practice.