Files
ytplayer/docs/p2p-architecture.md

10 KiB
Raw Permalink Blame History

Peer-to-peer video sharing — architecture (2026-09-29)

Source of truth for plans 008–019 in plans/queue/. Every P2P plan links here instead of repeating the rules. It replaces the earlier "phase 02–06" draft; the differences are listed at the end.

What the owner asked for

Peer-to-peer saving of videos. The server stores the user list, the video list and metadata. While the original source is online, a video is available for streaming and download. When it is downloaded to a device, the server records that device in the list of holders, so the video stays reachable from devices after the source is gone. Top or recent files stay on the server under a total space limit; files that don't meet the criteria (e.g. number of views, also stored on the server) are deleted first. The video id is the file hash of the highest-quality copy. The database grows over time. Each device has its own database that can be synced or added to the server's, with verification that the file exists. A file must first be downloaded by the server and checked before its hash is added to the server DB.

Corrections to the earlier draft (owner, 2026-09-29):

  1. P2P is ON by default (server and every client).
  2. The server's malware scan is OFF by default (admin opt-in). Hashing and the media validation gate are ALWAYS on and cannot be turned off.
  3. Holder records are persistent, not short-lived leases. A device stays listed as a holder until it says the file is gone, fails a check, or its reported list no longer contains it. The UI shows when each holder was last verified and marks it stale when that is older than P2P_STALE_DAYS (default 7). Online-right-now is a separate, live signal.

Vocabulary

Term Meaning
source Where bytes originally come from: YouTube (via yt-dlp) or a server upload (upl_…).
video id Existing ids (dQw4w9WgXcQ, upl_…). Still used everywhere in the app and API.
content id (cid) Lowercase hex SHA-256 of the exact file bytes. The P2P identity of a file. One video id can have several cids over time (a better master, the HEVC copy).
master The best copy the server keeps for a video: today the validated ≤720p H.264+AAC faststart MP4 in MEDIA_DIR (<id>.<gen>.mp4). The compression lane's HEVC copy is a second, separately hashed file. Raising the master quality later just creates new cids linked by video_id.
verified content A p2p_content row. Exists only after the server itself held the complete bytes, computed the SHA-256 itself and validateMedia() passed (plus the malware scan if enabled).
holder A device that reported holding a cid. Row in p2p_holders, never deleted by time.
online The device has an open /ws/p2p socket right now (in memory only).
stale now - last_verified_at > P2P_STALE_DAYS. Shown in the UI, still listed.

Data model (server, libsql — grows forever)

Added by plan 008 in server/p2p-db.js (initP2pSchema() runs after initDb()):

p2p_content (cid PK, video_id, size, height, vcodec, acodec, duration, meta JSON,
             origin 'server'|'intake', status 'verified'|'revoked',
             scan 'skipped'|'clean', created_at ms, verified_at ms)
p2p_devices (device_id PK 'dev_<16hex>', secret_hash, fingerprint, profile,
             share 0|1, created_at ms, last_seen_at ms)
p2p_holders (cid, device_id, status 'active'|'removed', trust 'reported'|'challenged',
             first_reported_at ms, last_verified_at ms, removed_at ms NULL,
             PRIMARY KEY (cid, device_id))
video_views (video_id, day 'YYYY-MM-DD', n, PRIMARY KEY (video_id, day))
media_cache.sha256  -- new column: cid of the current <id>.<gen>.mp4

"User list" = the existing users (fingerprints) and profiles tables plus p2p_devices. Nothing is ever deleted from p2p_content; a bad file is revoked.

Configuration (server/p2p-config.js, plan 008)

Env Default Meaning
P2P_ENABLED 1 (on) 0 turns off every P2P route, the hub and client features.
P2P_MALWARE_SCAN 0 (off) 1 runs P2P_SCAN_CMD <file> before admission; exit 0 = clean, 1 = infected (rejected), other = error (not admitted, retried later).
P2P_SCAN_CMD clamscan --no-summary --infected Needs an image built with --build-arg INSTALL_CLAMAV=1.
P2P_STALE_DAYS 7 Holder older than this is shown as stale.
P2P_KEEP_MIN_VIEWS 3 Retention: views in the last P2P_KEEP_DAYS that make a server copy "top".
P2P_KEEP_DAYS 30 Window for counting views.
P2P_KEEP_RECENT_DAYS 14 Retention: played this recently = "recent".
P2P_INTAKE_DIR <DB dir>/p2p-intake Quarantine for device uploads. Never served.
P2P_INTAKE_MAX_BYTES 3 GiB Largest accepted intake upload.

Client settings (data.settings, per profile): p2pShare: true (let other devices download my saved videos, and report holdings), p2pReceive: true (fetch from other devices when YouTube and the server can't serve).

Flows

  1. Server fetch (existing media cache) → verified content (plan 009). runFetch / runOptimize hash the promoted file, store media_cache.sha256, run the scan if enabled, then upsert p2p_content (origin 'server'). /api/download sends X-Content-SHA256. A backfill hashes already-cached files at boot, one at a time.
  2. Device save (plan 012). The OPFS worker hashes while it writes. If the server sent X-Content-SHA256 and the hash differs, the save fails (bonus integrity check). The device DB (IndexedDB ytp-device, store files) records {videoId, cid, size, savedAt, lastCheckedAt, state}; state is verified when the hashes matched, unverified when the server sent no hash, unhashed for old saves and the main-thread fallback path.
  3. Holdings sync (plan 013). Device registers once (POST /api/p2p/device → deviceId + secret, kept in localStorage.ytpDevice). It reports its holdings (POST /api/p2p/holdings, full list at launch, deltas after save/delete). The server accepts only cids in p2p_content with status='verified'; unknown cids come back in unknown (candidates for intake). When the server still has the file it returns up to 5 range challenges; a correct answer sets trust='challenged', a wrong one removes the holder. last_verified_at = time of the last report where the device re-checked the file (exists, same size; full re-hash every 30 days). A full report marks every active holder row of that device that is missing from the list as removed. There is no TTL.
  4. Presence (plan 014). /ws/p2p socket per device, authenticated with the device secret. Online status lives only in memory. The hub also relays WebRTC signalling between two online devices and carries server → device requests (plan 018).
  5. Availability (plans 014/015). GET /api/p2p/holders?v=<videoId> lists cids and their holders: opaque peer id (never the fingerprint/profile), online, lastVerifiedAt, stale, trust, plus serverHas. The UI shows e.g. "📡 On 3 devices · 1 online now · last checked 2 d ago", with stale holders greyed.
  6. Peer download (plan 017). WebRTC data channel (STUN only, same ICE list as watch party), 64 KiB frames with bufferedAmount back-pressure, receiver writes through a worker into OPFS while hashing; only a matching SHA-256 is committed. The new copy is a holder at the next report. Download-then-play; no progressive peer streaming. Used when /api/streams fails and the device has no copy, and from a "Get from a device" button.
  7. Intake (plan 016). A device can hand a file to the server (POST /api/p2p/intake → ticket, PUT the bytes). The server writes it to the quarantine dir, hashes it, runs validateMedia(), runs the scan if enabled, and only then inserts p2p_content (origin 'intake'). If the server has no copy of that video it adopts the file into the media cache (budget permitting).
  8. Rehydrate (plan 018). When a video is requested, its source fails, the server evicted its copy, and a verified holder is online with p2pShare on, the hub asks that device to upload it through intake (known cid → quick accept).
  9. Retention (plan 010). Views are counted per video per day. When the media cache needs room it evicts in this order: copies that are neither "top" (views in P2P_KEEP_DAYS ≥ P2P_KEEP_MIN_VIEWS) nor "recent" (played within P2P_KEEP_RECENT_DAYS), fewest views first, then oldest; only then the qualifying ones by LRU. The 10-minute play protection and MEDIA_CACHE_MAX_BYTES stay. Evicting a server copy never deletes p2p_content or holder rows.

Security rules every plan must keep

  • No cid enters p2p_content unless the SERVER computed it over bytes it holds and validateMedia passed. Clients can never insert or edit content rows, views or trust.
  • Intake files live in P2P_INTAKE_DIR, never under ./public or MEDIA_DIR, and are deleted on failure.
  • Device secrets: 32 random bytes, only sha256(secret) stored, compared with timingSafeEqual.
  • Holder lists never expose fingerprints, profile names or IPs; a peer id is sha256('peer:' + device_id).slice(0, 12).
  • The hub relays signalling only between two authenticated, online devices, with a per-socket message budget.
  • P2P_ENABLED=0 must leave the rest of the app working exactly as before.
  • Jobs stay server-owned; never pass a request AbortSignal into them (CLAUDE.md).

Where this differs from the earlier phase 02–06 draft

Earlier draft Now
"Default-off" P2P subsystem; "no inventory/upload from default settings" P2P on by default; devices report holdings and seed by default (can be turned off).
Mandatory scanner, "scan skip is failure" Scanner off by default (P2P_MALWARE_SCAN=0); hash + validateMedia mandatory.
Short-lived online leases; "cache availability only as an expiring hint" Persistent holder rows with last_verified_at and a stale marker; online status is separate.
Migrate all localStorage (_ytpdata) to IndexedDB first Not now: the device DB holds files + cids only. _ytpdata stays in localStorage (lower risk).
Collections, invitations, scoped principals, signed manifests, TURN, renditions lineage Deferred. Scope is one shared catalog + device secrets; add later if needed.
Separate server/p2p/* directory with migrations ledger Flat files server/p2p-*.js matching the repo's style (party.js, remote.js).