60 lines
7.2 KiB
Markdown
60 lines
7.2 KiB
Markdown
# Recommendations and the video catalog
|
||
|
||
Home and the Search menu show **Recommended for you**, with play, queue, playlist actions, refresh, and an explanation per pick. Saved picks remain available when the server cannot be reached. Cold starts explain how to get recommendations without blocking search.
|
||
|
||
## Server collection
|
||
|
||
`video-catalog.js` owns collection independently of downloaded media and the search-result cache. It preserves richer fields when sparse cards arrive and serializes writes so simultaneous discoveries cannot erase each other's metadata.
|
||
|
||
| Source | Collection path |
|
||
| --- | --- |
|
||
| Fresh YouTube searches | Awaited ingestion in `fetchYoutube`, including background refreshes; Innertube cards or yt-dlp fallback cards |
|
||
| Memory/persistent search hits | Successful JSON response collector |
|
||
| Browser saved searches | `/api/catalog/collect`, including the IndexedDB-only fast path; pending query keys retry on launch, reconnect, and every minute |
|
||
| Channels, expanded/shared playlists, related-video searches | Successful JSON response collector |
|
||
| Played, warmed, or saved videos | Full yt-dlp extraction in `resolveStreamsUncached`, plus cached-stream responses |
|
||
| Local/imported playlists, queue, history, linked profiles | Successful sync/profile/share requests and read responses |
|
||
| Other GET APIs surfacing video cards | Shared successful-JSON response collector; recommendations themselves are excluded |
|
||
| Older server data | Durable paged backfill of known cards, search results, media metadata, history, playlists, shared playlists, and profiles |
|
||
|
||
The catalog records YouTube video IDs, titles, channels, duration, thumbnail URL, available tags/categories, description excerpts, and extractor view counts. It tracks discovery source separately from listening counts. Local uploads already have server metadata/art in the upload tables; custom device-only edits are excluded from the public YouTube catalog.
|
||
|
||
Thumbnail **bytes** are stored in SQLite, not merely URLs. Two workers fetch allowed HTTPS YouTube image hosts, reject redirects, allow JPEG/PNG/WebP, cap images at 1 MiB, and time out after eight seconds. Failed jobs persist with backoff up to one day and resume after restart. `/api/catalog/:id/thumbnail` serves stored images; recommendation and local-search cards prefer this URL once ready. The service worker caches these URLs for offline images.
|
||
|
||
Retention is bounded: `VIDEO_META_MAX` defaults to 500,000 videos and `VIDEO_THUMB_MAX_BYTES` to 512 MiB of image data. Metadata trimming prefers unplayed entries and cleans up associated jobs/source/channel rows. Image eviction preserves metadata, disables automatic re-download, and permits re-fetch on a new discovery. SQLite can reuse freed pages; the byte budget measures live image data, not the physical database file. Failed or evicted images fall back to the original thumbnail URL. Existing source images can be unavailable; metadata collection does not require successful image retrieval.
|
||
|
||
## Listening and ranking
|
||
|
||
The existing StatsCore tracker counts a play after **30 seconds actually listened**, rather than counting search appearances, stream requests, or media-cache hits. `/api/user/sync` submits daily per-video snapshots. The server uses monotonic maximum counts per listener/day/video so repeated saves, retries, and reloads do not multiply plays. It retains 400 days and caps ingestion at 10,000 day/video entries per sync.
|
||
|
||
Anonymous listeners use device fingerprints. Linked listeners use the profile name and existing profile password/key gate. Linking devices merges earlier anonymous daily summaries into the profile with maxima and removes the device copies, so synchronized profile history is counted once. Recommendations and browser snapshots switch with the active profile. This uses the app's cumulative daily statistics; it does not introduce event-level reconciliation of independently edited or concurrently modified profile histories.
|
||
|
||
The algorithm runs on the server in `recommendations.js`:
|
||
|
||
1. Select up to 300 globally most-played videos, 100 personal favorites, and 600 recent discoveries. Fetch older candidates from up to ten favorite channels using an indexed channel table.
|
||
2. Build channel and title/tag/category interests from the listener's top 20 videos.
|
||
3. Score with logarithmic weights for personal plays, channel affinity, shared metadata tokens, global plays, and plays in the last 30 days.
|
||
4. Keep at most three videos from a channel and reserve up to half the picks for familiar favorites when discoveries exist.
|
||
5. Return playable cards and plain-language reasons; prefer stored thumbnail URLs.
|
||
|
||
There is no `media_cache` filter or dependency. Unplayed search results and cached or uncached videos all qualify. An empty listener history falls back to global most-played videos, then recent discoveries. A completely empty catalog returns an empty list; recommendations do not launch unsolicited extraction searches.
|
||
|
||
## API
|
||
|
||
- `GET /api/recommendations?fp=<fingerprint>[&profile=<name>][&limit=12]`: `{ok, results}`; limit 1–24. Protected profiles require `X-Profile-Secret`. Personal responses use `Cache-Control: no-store`.
|
||
- `POST /api/catalog/collect`: `{cards: [...]}`, at most 500 cards and 512,000 request characters. Normalized cards reject unsafe IDs and arbitrary thumbnail hosts.
|
||
- `GET /api/catalog/:id/thumbnail`: stored image bytes or 404.
|
||
- `/api/user/sync` additionally accepts `stats` and `profileName`; protected profiles use the existing secret header.
|
||
|
||
## Verification
|
||
|
||
- 16 Bun catalog/recommendation tests: all pass, covering source collection, richer-field preservation, concurrency, cache eviction, thumbnails, durable retry/backfill, ranking, indexed retrieval of older channel matches, idempotent analytics, profile scope and deduplication.
|
||
- 9 Chromium recommendation UI tests: all pass, covering 320/390/1440px, caption contrast >=4.5:1, actions, Search navigation during playback, saved-search submission, offline retries, failure states and profile switching.
|
||
- Existing 11 Classic browser checks and 72 frontend unit tests: all pass.
|
||
- Isolated live Bun server: protected-profile access, recommendation persistence across restart, two-device deduplication (eight plays remain eight), and a real 21,011-byte YouTube thumbnail stored and served after restart all pass. No production profiles or media were changed.
|
||
- Visual review: desktop cards and both mobile section captures are readable and reachable. Initial mobile captures were taken after scrolling to the final action; corrected captures show the section heading and first cards. Desktop retained visible Try chips because its content did not need to scroll; this is a capture expectation difference, not a clipping defect.
|
||
|
||
Design-hook triage: the new recommendation stylesheet has no findings and measured Classic caption contrast passes. One file/value exception suppresses the shared HTML's three intentionally empty, hidden dynamic artwork images; JavaScript supplies their `src`. The other 15 inherited shared-shell findings (legacy contrast, gradient/glow styling, 11px captions, and intentional app scroll-container clipping) remain unsuppressed and outside this feature's styling changes.
|
||
|
||
No real mobile device/WebKit run or production deployment was performed.
|