Resumable saves: ranged downloads that survive dropped connections, reloads and app kills

Server
- /api/download/:id answers Range with a strong ETag ("<id>.<gen>") and honours
  If-Range; a stale partial gets the whole current file (streamed, so Bun does
  not re-apply the Range itself). Uploads get the same treatment.
- GET /api/download/:id/prepare never blocks: ready {gen,size,sha256,etag,ext},
  working (server still fetching), legacy (HEVC / too long / cache offline),
  failed. Recently refused prepares are remembered for 10 minutes.

Browser
- The OPFS worker saves in 8 MiB ranges, writes at the byte offset, flushes
  each chunk, retries each chunk 6 times with backoff (30 s idle timeout) and
  keeps the .part plus a .part.json sidecar naming the server copy it belongs
  to. A changed copy restarts cleanly; the finished file is hashed once and
  checked against the server's SHA-256.
- SaveQueue remembers unfinished saves and resumes them on start, online,
  return to the foreground and every 2 minutes while visible; one at a time.
- Downloads shows live MB progress, 'Preparing on server', 'Verifying', and
  paused saves with Resume and Cancel (confirmed).
- listVideos ignores the sidecars; new listPartials/discardPartial helpers.

Verified in Chromium through a connection-dropping proxy: paused at 8 MiB,
auto-resumed after a reload from byte 8388608, final SHA-256 matched.
This commit is contained in:
Jonathan Sykes
2026-10-03 00:46:52 +08:00
parent c8601c5953
commit 56b3f4e449
10 changed files with 698 additions and 63 deletions

View File

@@ -0,0 +1,138 @@
# Resumable save / download — plan
Status: **implemented 2026-10-03** (items 1–7; item 6 without the wake lock).
Verified end to end in Chromium through a connection-dropping proxy: a 19 MB
save paused after its retries with 8 MiB kept, resumed by itself after a
reload from byte 8388608, and the stored file's SHA-256 matched the server's.
As built: prepare is `GET /api/download/:id/prepare` (states ready / working /
legacy / failed); the partial's owner is a `<file>.part.json` sidecar in OPFS
(no IndexedDB record); the queue of unfinished saves is `localStorage.ytpSaveQueue`
(`SaveQueue` in app.js); pure helpers live in `frontend/resume-core.js`. Goal: saving a long video (≥ 1 h, hundreds of MB)
to the device must survive dropped connections, app backgrounding, reloads and
the slow homelab link — continuing from where it stopped instead of starting over.
## Why long saves get stuck today
| Piece | Today | Problem for big files |
|---|---|---|
| `frontend/app.js` `opfsDownload()` (~273) | one `fetch('/api/download/<id>')` for the whole file | any drop = total loss; no progress survives a reload |
| `frontend/opfs.js` `writeFromResponse()` (~178), `opfs-worker.js` (~40) | writes a `.part`, checks `Content-Length`, **deletes the `.part` on any failure** | the bytes already received are thrown away |
| `server/server.js` `/api/download/:id` (~1549) → `cachedDownloadResponse()` (~1527) | streams the cached copy with `Content-Length` only | **no `Range` / `Accept-Ranges`** → the client cannot ask for "the rest" |
| `/api/download/:id` when the server has no copy yet (~1620) | streams yt-dlp output live | not resumable at all; one stall kills it |
| iOS / Safari | backgrounding suspends `fetch` | the single long request dies whenever the phone locks |
Already in place and reusable: `rangeFileResponse()` (~1265, used by
`/api/media/:id` and `/api/export/:id`) does correct 206/`Content-Range`;
`/api/media/:id/status` reports cache-job progress; `sha256.js` hashes
incrementally; the OPFS worker uses `createSyncAccessHandle`, which can write at
any offset; Bun `idleTimeout` is already 0.
## Design
**Rule: the device only ever downloads a finished, validated server copy, in
byte ranges.** Fetching from YouTube is the server's job (the media cache),
never part of a device transfer.
```
tap Save ──► POST /api/media/:id/prepare ──► server media-cache job (HIGH)
│ │ status/progress
│◄── poll GET /api/media/:id/status ◄────────┘ "Preparing on server 42%"
▼ ready: { gen, size, sha256 }
for each missing chunk (8 MiB):
GET /api/media/:id?g=<gen> Range: bytes=a-b If-Range: "<id>.<gen>"
write at offset a (OPFS sync handle, in the worker)
record bytesDone in IndexedDB
all chunks ──► hash whole .part, compare sha256 ──► rename to final
```
### Server (Bun) — `server/server.js`, `server/media-cache.js`
1. **Range on the download path.** `/api/download/:id` with a cached copy (and
uploads) answers through `rangeFileResponse()`, plus `Accept-Ranges: bytes`
and a strong `ETag: "<id>.<gen>"` (uploads: `"<id>"`). Honor `If-Range`: when
the ETag no longer matches (copy was re-downloaded → new gen), send a full 200
so the client knows to restart. Keep `X-Content-SHA256`.
2. **Generation-pinned URLs.** Chunks use `/api/media/:id?g=<gen>` (already
immutable + ranged). A finished copy is never rewritten in place (new gen =
new file), so bytes for one gen never change under the client.
3. **Pin while downloading.** Eviction must not remove a copy a device is in the
middle of fetching: `media.touch(id)` on every ranged hit already refreshes
LRU; add a short "in transfer" protect window (reuse `EVICT_PROTECT_MS`).
4. **Prepare endpoint.** `POST /api/media/:id/prepare` → `ensureCached(id, {priority: HIGH})`
and return `status()` immediately (no waiting). `status()` gains
`{ progress, gen, size, sha256 }` so the client can show
"Preparing on server" with a percentage and knows the exact target.
5. **Long videos.** `MAX_SAVE_SECONDS` / `MEDIA_AUTO_MAX_SECONDS` stay as the
policy limits; when they refuse, `prepare` returns the reason so the UI can
say so instead of hanging. The yt-dlp live-stream path of `/api/download`
stays only as a legacy fallback for short videos.
6. **Uploads.** Same Range/ETag treatment (they already sit on disk; with the
USB drive move, the backup-dir fallback in `uploads.js` serves either copy —
both are byte-identical, so the ETag is the same).
### Browser — `frontend/app.js`, `frontend/opfs-worker.js`, `frontend/opfs.js`, `frontend/device-db.js`
1. **Transfer record (IndexedDB, `device-db.js`).** One row per saving video:
`{ id, gen, size, sha256, chunk, bytesDone, state: preparing|downloading|verifying|done|failed, updatedAt, error }`.
This is what survives reloads, crashes and app kills.
2. **Chunked worker download (`opfs-worker.js`).** Replace the single fetch with a
loop over missing ranges: `Range: bytes=start-end`, `If-Range: "<id>.<gen>"`,
per-chunk `AbortController` timeout (60 s), up to 5 retries with backoff
(2 s → 30 s). Write with `accessHandle.write(buf, { at: start })`, `flush()`
every chunk, update `bytesDone` after the flush (so the record never claims
bytes that are not on disk). On start, trust `min(record.bytesDone, .part size)`.
- 200 instead of 206 or an ETag mismatch → the server copy changed: truncate
the `.part`, reset the record to the new gen, start over (rare).
- 416 → `.part` is longer than the file: truncate to `size` and verify.
3. **Never delete progress on failure.** `.part` + record are kept on errors and
only removed on user cancel, on success, or when the record is older than 7
days (sweep at startup, alongside OPFS quota checks).
4. **Verify at the end, not during.** Hashing across resumes is done by
re-reading the finished `.part` in the worker with `sha256.js` (no hash state
to persist), then compare with the server's `sha256`; mismatch → discard and
restart once, then mark failed. Rename `.part` → final as today.
5. **Auto-resume triggers.** App start, `online`, `visibilitychange` → visible,
and the SW `sync` event where supported: resume every record in
`downloading`/`preparing`. One transfer at a time (the homelab uplink is the
bottleneck), next in queue starts when one finishes.
6. **Browsers without worker sync access handles.** Detect up front. Where only
`createWritable` exists (some desktop Chromium contexts), use
`createWritable({ keepExistingData: true })` + `seek(start)` per chunk. If
neither exists, keep today's "not supported" error — do not fall back to an
unresumable path silently.
7. **UI.**
- Save button / Downloads list show `Preparing on server 40%`,
`Downloading 312 / 742 MB`, `Paused — will resume`, `Verifying…`.
- Pause / Resume / Cancel per item in **Downloads** (Cancel deletes `.part`).
- A paused item resumes by itself on the triggers above; a toast only on
final success or a hard failure (with the reason from the server).
- Keep the screen awake (existing wake-lock helper) while a foreground
download runs, opt-out in Settings, so iOS does not suspend it.
8. **Save-to-device export** (`exportToDevice`, ~8000) already uses the ranged
`/api/export/:id`; no change beyond using the same ETag.
### Browser support matrix (target)
| Browser | Write at offset | Resume across reload | Notes |
|---|---|---|---|
| Chrome / Edge / Android Chrome | worker sync handle | yes | |
| Safari / iOS 16.4+ (PWA + tab) | worker sync handle | yes | suspended when backgrounded → resumes on `visibilitychange` |
| Firefox 111+ | worker sync handle | yes | |
| Older browsers without OPFS sync handles | — | — | clear "not supported" message, Save-to-device export still works |
## Work items (each one commit, tests first)
1. Server: Range + strong ETag + `If-Range` on `/api/download/:id` (cached + uploads). Tests: 206 slices, full 200 on ETag mismatch, 416 past the end.
2. Server: `POST /api/media/:id/prepare`; `status()` returns `{ progress, gen, size, sha256 }`; in-transfer eviction protect. Tests in `media-cache.test.js`.
3. Browser: transfer record store in `device-db.js` + startup sweep (node tests with a fake IDB).
4. Browser: chunked, resumable worker download in `opfs-worker.js` (pure chunk planner + retry policy as testable functions; node tests with a fake fetch that drops mid-chunk, returns 200 on ETag change, and 416).
5. Browser: wire `opfsDownload()` / `preload()` to prepare → poll → chunked download; auto-resume triggers; one-at-a-time queue.
6. UI: progress states, Pause / Resume / Cancel in Downloads, wake lock while downloading.
7. End-to-end check against a local server with a 1 h+ fixture: kill the network mid-way (Playwright `context.setOffline`), reload the page, confirm it resumes from the last chunk and the final SHA-256 matches; run the same in WebKit (Windows Playwright, see CLAUDE.md) for Safari behaviour.
## Out of scope
- Resuming the server-side YouTube fetch itself (the media cache already restarts
jobs on boot and retries with backoff).
- Background downloads while the iOS app is fully closed (no Background Fetch on
iOS); the transfer resumes the next time the app is opened.