Files
ytplayer/docs/resumable-downloads-plan.md
Jonathan Sykes 56b3f4e449 Resumable saves: ranged downloads that survive dropped connections, reloads and app kills
Server
- /api/download/:id answers Range with a strong ETag ("<id>.<gen>") and honours
  If-Range; a stale partial gets the whole current file (streamed, so Bun does
  not re-apply the Range itself). Uploads get the same treatment.
- GET /api/download/:id/prepare never blocks: ready {gen,size,sha256,etag,ext},
  working (server still fetching), legacy (HEVC / too long / cache offline),
  failed. Recently refused prepares are remembered for 10 minutes.

Browser
- The OPFS worker saves in 8 MiB ranges, writes at the byte offset, flushes
  each chunk, retries each chunk 6 times with backoff (30 s idle timeout) and
  keeps the .part plus a .part.json sidecar naming the server copy it belongs
  to. A changed copy restarts cleanly; the finished file is hashed once and
  checked against the server's SHA-256.
- SaveQueue remembers unfinished saves and resumes them on start, online,
  return to the foreground and every 2 minutes while visible; one at a time.
- Downloads shows live MB progress, 'Preparing on server', 'Verifying', and
  paused saves with Resume and Cancel (confirmed).
- listVideos ignores the sidecars; new listPartials/discardPartial helpers.

Verified in Chromium through a connection-dropping proxy: paused at 8 MiB,
auto-resumed after a reload from byte 8388608, final SHA-256 matched.
2026-10-03 00:46:52 +08:00

139 lines
9.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Resumable save / download — plan
Status: **implemented 2026-10-03** (items 1–7; item 6 without the wake lock).
Verified end to end in Chromium through a connection-dropping proxy: a 19 MB
save paused after its retries with 8 MiB kept, resumed by itself after a
reload from byte 8388608, and the stored file's SHA-256 matched the server's.
As built: prepare is `GET /api/download/:id/prepare` (states ready / working /
legacy / failed); the partial's owner is a `<file>.part.json` sidecar in OPFS
(no IndexedDB record); the queue of unfinished saves is `localStorage.ytpSaveQueue`
(`SaveQueue` in app.js); pure helpers live in `frontend/resume-core.js`. Goal: saving a long video (≥ 1 h, hundreds of MB)
to the device must survive dropped connections, app backgrounding, reloads and
the slow homelab link — continuing from where it stopped instead of starting over.
## Why long saves get stuck today
| Piece | Today | Problem for big files |
|---|---|---|
| `frontend/app.js` `opfsDownload()` (~273) | one `fetch('/api/download/<id>')` for the whole file | any drop = total loss; no progress survives a reload |
| `frontend/opfs.js` `writeFromResponse()` (~178), `opfs-worker.js` (~40) | writes a `.part`, checks `Content-Length`, **deletes the `.part` on any failure** | the bytes already received are thrown away |
| `server/server.js` `/api/download/:id` (~1549) → `cachedDownloadResponse()` (~1527) | streams the cached copy with `Content-Length` only | **no `Range` / `Accept-Ranges`** → the client cannot ask for "the rest" |
| `/api/download/:id` when the server has no copy yet (~1620) | streams yt-dlp output live | not resumable at all; one stall kills it |
| iOS / Safari | backgrounding suspends `fetch` | the single long request dies whenever the phone locks |
Already in place and reusable: `rangeFileResponse()` (~1265, used by
`/api/media/:id` and `/api/export/:id`) does correct 206/`Content-Range`;
`/api/media/:id/status` reports cache-job progress; `sha256.js` hashes
incrementally; the OPFS worker uses `createSyncAccessHandle`, which can write at
any offset; Bun `idleTimeout` is already 0.
## Design
**Rule: the device only ever downloads a finished, validated server copy, in
byte ranges.** Fetching from YouTube is the server's job (the media cache),
never part of a device transfer.
```
tap Save ──► POST /api/media/:id/prepare ──► server media-cache job (HIGH)
│ │ status/progress
│◄── poll GET /api/media/:id/status ◄────────┘ "Preparing on server 42%"
▼ ready: { gen, size, sha256 }
for each missing chunk (8 MiB):
GET /api/media/:id?g=<gen> Range: bytes=a-b If-Range: "<id>.<gen>"
write at offset a (OPFS sync handle, in the worker)
record bytesDone in IndexedDB
all chunks ──► hash whole .part, compare sha256 ──► rename to final
```
### Server (Bun) — `server/server.js`, `server/media-cache.js`
1. **Range on the download path.** `/api/download/:id` with a cached copy (and
uploads) answers through `rangeFileResponse()`, plus `Accept-Ranges: bytes`
and a strong `ETag: "<id>.<gen>"` (uploads: `"<id>"`). Honor `If-Range`: when
the ETag no longer matches (copy was re-downloaded → new gen), send a full 200
so the client knows to restart. Keep `X-Content-SHA256`.
2. **Generation-pinned URLs.** Chunks use `/api/media/:id?g=<gen>` (already
immutable + ranged). A finished copy is never rewritten in place (new gen =
new file), so bytes for one gen never change under the client.
3. **Pin while downloading.** Eviction must not remove a copy a device is in the
middle of fetching: `media.touch(id)` on every ranged hit already refreshes
LRU; add a short "in transfer" protect window (reuse `EVICT_PROTECT_MS`).
4. **Prepare endpoint.** `POST /api/media/:id/prepare` → `ensureCached(id, {priority: HIGH})`
and return `status()` immediately (no waiting). `status()` gains
`{ progress, gen, size, sha256 }` so the client can show
"Preparing on server" with a percentage and knows the exact target.
5. **Long videos.** `MAX_SAVE_SECONDS` / `MEDIA_AUTO_MAX_SECONDS` stay as the
policy limits; when they refuse, `prepare` returns the reason so the UI can
say so instead of hanging. The yt-dlp live-stream path of `/api/download`
stays only as a legacy fallback for short videos.
6. **Uploads.** Same Range/ETag treatment (they already sit on disk; with the
USB drive move, the backup-dir fallback in `uploads.js` serves either copy —
both are byte-identical, so the ETag is the same).
### Browser — `frontend/app.js`, `frontend/opfs-worker.js`, `frontend/opfs.js`, `frontend/device-db.js`
1. **Transfer record (IndexedDB, `device-db.js`).** One row per saving video:
`{ id, gen, size, sha256, chunk, bytesDone, state: preparing|downloading|verifying|done|failed, updatedAt, error }`.
This is what survives reloads, crashes and app kills.
2. **Chunked worker download (`opfs-worker.js`).** Replace the single fetch with a
loop over missing ranges: `Range: bytes=start-end`, `If-Range: "<id>.<gen>"`,
per-chunk `AbortController` timeout (60 s), up to 5 retries with backoff
(2 s → 30 s). Write with `accessHandle.write(buf, { at: start })`, `flush()`
every chunk, update `bytesDone` after the flush (so the record never claims
bytes that are not on disk). On start, trust `min(record.bytesDone, .part size)`.
- 200 instead of 206 or an ETag mismatch → the server copy changed: truncate
the `.part`, reset the record to the new gen, start over (rare).
- 416 → `.part` is longer than the file: truncate to `size` and verify.
3. **Never delete progress on failure.** `.part` + record are kept on errors and
only removed on user cancel, on success, or when the record is older than 7
days (sweep at startup, alongside OPFS quota checks).
4. **Verify at the end, not during.** Hashing across resumes is done by
re-reading the finished `.part` in the worker with `sha256.js` (no hash state
to persist), then compare with the server's `sha256`; mismatch → discard and
restart once, then mark failed. Rename `.part` → final as today.
5. **Auto-resume triggers.** App start, `online`, `visibilitychange` → visible,
and the SW `sync` event where supported: resume every record in
`downloading`/`preparing`. One transfer at a time (the homelab uplink is the
bottleneck), next in queue starts when one finishes.
6. **Browsers without worker sync access handles.** Detect up front. Where only
`createWritable` exists (some desktop Chromium contexts), use
`createWritable({ keepExistingData: true })` + `seek(start)` per chunk. If
neither exists, keep today's "not supported" error — do not fall back to an
unresumable path silently.
7. **UI.**
- Save button / Downloads list show `Preparing on server 40%`,
`Downloading 312 / 742 MB`, `Paused — will resume`, `Verifying…`.
- Pause / Resume / Cancel per item in **Downloads** (Cancel deletes `.part`).
- A paused item resumes by itself on the triggers above; a toast only on
final success or a hard failure (with the reason from the server).
- Keep the screen awake (existing wake-lock helper) while a foreground
download runs, opt-out in Settings, so iOS does not suspend it.
8. **Save-to-device export** (`exportToDevice`, ~8000) already uses the ranged
`/api/export/:id`; no change beyond using the same ETag.
### Browser support matrix (target)
| Browser | Write at offset | Resume across reload | Notes |
|---|---|---|---|
| Chrome / Edge / Android Chrome | worker sync handle | yes | |
| Safari / iOS 16.4+ (PWA + tab) | worker sync handle | yes | suspended when backgrounded → resumes on `visibilitychange` |
| Firefox 111+ | worker sync handle | yes | |
| Older browsers without OPFS sync handles | — | — | clear "not supported" message, Save-to-device export still works |
## Work items (each one commit, tests first)
1. Server: Range + strong ETag + `If-Range` on `/api/download/:id` (cached + uploads). Tests: 206 slices, full 200 on ETag mismatch, 416 past the end.
2. Server: `POST /api/media/:id/prepare`; `status()` returns `{ progress, gen, size, sha256 }`; in-transfer eviction protect. Tests in `media-cache.test.js`.
3. Browser: transfer record store in `device-db.js` + startup sweep (node tests with a fake IDB).
4. Browser: chunked, resumable worker download in `opfs-worker.js` (pure chunk planner + retry policy as testable functions; node tests with a fake fetch that drops mid-chunk, returns 200 on ETag change, and 416).
5. Browser: wire `opfsDownload()` / `preload()` to prepare → poll → chunked download; auto-resume triggers; one-at-a-time queue.
6. UI: progress states, Pause / Resume / Cancel in Downloads, wake lock while downloading.
7. End-to-end check against a local server with a 1 h+ fixture: kill the network mid-way (Playwright `context.setOffline`), reload the page, confirm it resumes from the last chunk and the final SHA-256 matches; run the same in WebKit (Windows Playwright, see CLAUDE.md) for Safari behaviour.
## Out of scope
- Resuming the server-side YouTube fetch itself (the media cache already restarts
jobs on boot and retries with backoff).
- Background downloads while the iOS app is fully closed (no Background Fetch on
iOS); the transfer resumes the next time the app is opened.