Files
ytplayer/.agents/skills/lyrics-regenerate/SKILL.md

99 lines
4.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
name: lyrics-regenerate
description: Back up and replace lyrics on worship.hesed.sbs that are wrong, mistimed or misheard — typically machine transcripts — with published lyrics from LRCLIB. Use when the user says lyrics are incorrect/off/mistimed, names songs whose lyrics are bad, or asks to re-check songs on LRCLIB. Always backs up first and never writes an unverified artist match.
---
# Replace bad lyrics from LRCLIB
Machine transcripts (Whisper, Scribe) mishear words and drift out of time.
LRCLIB's published lyrics are usually correct and often **synced**. This skill
swaps them in — backup first, artist verified, one song at a time.
Script: `scripts/lyrics/lrclib_regen.py`.
## The rule: back up before you touch anything
The server keeps every previous version as a revision, and `/admin` can restore
one, but the script **also** writes an offline JSON backup of the current lyrics
of every song it will consider — on a dry run too. Do not skip it, do not write
your own one-off loop that lacks it.
```
/mnt/c/Users/josh/Documents/ytplayer-lyrics-backup-<YYYYmmdd-HHMMSS>.json
```
(`--backup-dir` or `YTP_BACKUP_DIR` to move it; on the devbox it falls back to `~/`.)
## Run it
```bash
cd ~/development/personal/ytplayer
export YTP_ADMIN_PASSWORD='…' # or YTP_TOKEN=ytp_…
# 1) ALWAYS dry-run first and read every line of the output
python3 scripts/lyrics/lrclib_regen.py --tagged auto-transcribed
# 2) apply once the matches look right
python3 scripts/lyrics/lrclib_regen.py --tagged auto-transcribed --apply
```
Picking the songs:
| Flag | Picks |
|---|---|
| `--ids A,B,C` | exactly those videos |
| `--tagged auto-transcribed` | songs whose lyrics carry that tag (machine transcripts) |
| `--since '2026-09-19 01:45' --until '2026-09-19 02:00'` | songs whose lyrics were **saved** in that window |
| *(none)* | every song the server has lyrics for |
`--since/--until` is the one to reach for when the user says *"the songs that got
lyrics at the same time as X"* — read X's `updatedAt` from
`GET /api/notes/<id>` and bracket it by a few minutes.
Other flags: `--tolerance 6` (max duration difference, seconds), `--loose`
(accept matches whose artist doesn't line up — risky, see below).
## Reading the output
```
eJBlOV6cM7Y LRCLIB synced | Israel Houghton – Holy You Are | 41 lines (was 38)
_n6dfB2Z-Ko UNSURE plain | Night Ranger – Still | 52 lines (was 44)
↳ artist doesn't match "Hillsong Worship" / the video title — left alone (use --loose to accept)
tYM05iaVu3I no match | Jesus At The Centre | … | keeping 60 lines (auto-transcribed)
```
- **LRCLIB** — verified match, will be written on `--apply`.
- **UNSURE** — title and duration fit but the artist doesn't appear in the channel
name or the video title. **Left alone by default. Do not pass `--loose` to make
it go away** — check the song by hand instead; this guard is what stopped a
Hillsong song being overwritten with a Night Ranger one.
- **no match** — LRCLIB doesn't have it. Existing lyrics are kept. Fall back to
`web_lyrics.py --agy --ids <id> --overwrite`, or fix it in the admin lyric editor.
Replaced songs are tagged `from LRCLIB (synced)` / `(plain text)`, which is also
how you tell later what has already been fixed.
## Restoring
- Per song, in the UI: `/admin` → recent edits → **Restore** on the older revision.
- From the JSON backup: `PUT /api/notes/<id>/lyrics` with
`{"data": <songs[id].data>, "baseRev": <current rev from GET /api/notes/<id>>}`.
Use the *current* rev, not the backed-up one — `baseRev` is optimistic
concurrency, not a version to travel back to.
## Gotchas
- **LRCLIB rate-limits**: 503/429 on bursts. `http_json()` retries with backoff and
the loop sleeps 0.4 s between songs. A song that fails all retries is reported and
skipped — re-run it later rather than hammering.
- **Duration match is ±6 s** against `/api/streams` metadata. Live or extended cuts
legitimately miss; raise `--tolerance` deliberately, per song.
- **Title cleaning** strips "(Official Video)", "Lyrics", "[HD]" etc.
`clean_title`'s `NOISE` regex is used with `.sub()` and `.search()` — never give
it a `/g`-style shared match state; a stateful regex silently skipped every other
song once already.
- **Karaoke/minus-one** versions match the original recording's lyrics, which is
usually what you want, but the timing won't line up. Check before applying.
Related skills: `lyrics-lookup` (songs with **no** lyrics), `deploy-prod`.