Files
ytplayer/.agents/skills/lyrics-agy/SKILL.md

154 lines
6.8 KiB
Markdown

---
name: lyrics-agy
description: Find lyrics (with timings when they exist) for worship.hesed.sbs songs using agy, the flat-rate Antigravity CLI, via the agy-bridge MCP server. Use when LRCLIB has no match, when a song still has no lyrics or bad machine-transcribed ones, or when the user says "ask agy for the lyrics". Also how to paste lyrics you already have into a song.
---
# Lyrics from agy
agy searches the open web and, when the song has a caption track or a published
sync, returns **timed** lyrics in LRC form. It is flat-rate, so a run costs
nothing per song — but it is the **last** resort, after LRCLIB (`lyrics-lookup`,
`lyrics-regenerate`), because LRCLIB's synced lyrics are published data while
agy's are derived.
Script: `scripts/lyrics/agy_lyrics.py`.
## Run it
```bash
cd ~/development/personal/ytplayer
export YTP_ADMIN_PASSWORD='…' # or YTP_TOKEN=ytp_…
# every song that still has no lyrics — dry run first, ALWAYS
python3 scripts/lyrics/agy_lyrics.py --missing
python3 scripts/lyrics/agy_lyrics.py --missing --apply
# one song, replacing lyrics that are wrong
python3 scripts/lyrics/agy_lyrics.py --ids ZHl6EwSwjv0 --overwrite --apply
# lyrics you already have (LRC or plain text), no agy call at all
python3 scripts/lyrics/agy_lyrics.py --ids ZHl6EwSwjv0 --from-file words.txt --overwrite --apply
```
Every run backs the current lyrics up to
`Documents/ytplayer-lyrics-backup-<stamp>.json` before writing, and the server
keeps each previous version as a restorable revision.
**One song takes ~2 minutes per attempt**, and `--tries` defaults to 3, so a
batch is slow. Run it detached rather than in a foreground command that will hit
a timeout:
```bash
setsid nohup python3 scripts/lyrics/agy_lyrics.py --missing --apply \
> /tmp/agy-lyrics.log 2>&1 < /dev/null &
```
## How it calls agy
Through MCPJungle, so no agy CLI contract is hard-coded here:
```bash
mcpjungle invoke agy-bridge__agy_ask --input '{"dir": "<repo>", "prompt": "…"}'
mcpjungle invoke agy-bridge__fetch_output --input '{"keep_id": "ask-…"}'
```
Four things about that interface cost real debugging time — do not rediscover them:
1. **`dir` is required.** Without it the call fails schema validation.
2. **The answer arrives on STDERR**, not stdout. Read both streams or you get
an empty string and conclude, wrongly, that agy found nothing.
3. **Long answers are truncated** with `fetch_output(keep_id='…') for more`.
A full set of lyrics is almost always longer than the cap, so always follow
the `keep_id` — otherwise you silently save half a song.
4. **agy can fail and still look like a success.** A quota-exhausted instance
returns prose like `WHY: agy-ask failed (rc=1)` inside a `STATUS: ok`
envelope. One such line even carried a `[05:37.76]` stamp and parsed as a
perfectly good synced lyric. `FAILED` in the script rejects those; keep it.
## It is not deterministic — that is the main gotcha
The same question can come back synced, plain, or empty on consecutive calls.
Observed in one sitting on "Jesus At The Centre": 39 timed lines, then 43
untimed, then nothing. So the script asks up to `--tries` times and keeps the
**best** answer (`score()`: any timed lines beats none, then more lines beats
fewer), stopping early once a timed answer arrives.
If a song saves untimed and you believe a timed version exists, just run it
again with `--overwrite`.
## What it filters out, and why
agy streams its own progress into the answer. Everything below is dropped
before parsing, and every pattern is there because it once ended up saved as a
lyric line:
- `STATUS:` / `SUMMARY:` / `MODEL:` / `INSTANCE:` / `AGY-META:` / `WHY:` envelopes
- `Waiting for task execution…`, `Background task <uuid> completed with…`
- section labels that are not sung (`Verse 1`, `[Chorus]`, `x2`)
- anything over 200 characters (a paragraph of commentary, not a sung line)
After filtering, an answer shorter than `--min-lines` (6) is rejected outright.
## Reviewing before you trust it
agy derives timings, so check them once per song before relying on them in a
service:
- The last timestamp should land near the song's length (a 384 s song ending at
367 s is right; one ending at 120 s means it only got a verse).
- Timestamps must increase monotonically.
- Open `/admin?v=<id>`, press play and watch the highlighted line track the
vocal. Fix drift with **Shift all**, or retime individual lines with **Set** /
**Tap mode**.
Untimed results are fine — the app shows them as a plain scrolling list, and
Tap mode turns them into synced lyrics in one pass of the song.
## Provenance tags
Written into `data.tags` so a later run can tell where lyrics came from:
| Tag | Meaning |
|---|---|
| `from the web via agy (synced)` | agy, with timings |
| `from the web via agy (untimed) — check and Tap-sync` | agy, words only |
| `from a file (<name>) (synced\|untimed)` | `--from-file` |
## Right words + real timings: `retime_lyrics.py`
The two sources fail in opposite ways — the web (LRCLIB plain, agy) has the
right words and line breaks but no timings; Whisper has a time for every line
but mishears words and breaks lines mid-phrase ("To show for the / years").
Rather than re-rolling agy for a timed answer that may never come, align them:
```bash
# correct words in a file (LRCLIB plain text, agy output, or pasted lyrics)
python3 scripts/lyrics/retime_lyrics.py --id MU5dlCRTLY8 --words correct.txt
python3 scripts/lyrics/retime_lyrics.py --id MU5dlCRTLY8 --words correct.txt --apply
# take timings from a backup rather than what is live
… --id ID --words w.txt --times-from ytplayer-lyrics-backup-….json
```
It matches on a **word stream, not line to line**, precisely because the line
breaks disagree: each transcript line's words are spread across the gap to the
next line, `difflib` aligns the two word sequences, and a correct line takes the
time of the earliest transcript word inside it. Times are forced monotonic, and
lines that matched nothing are interpolated between their neighbours. On "Take
Me to the End": 46 correct lines, 44 timed directly, 2 interpolated — and the
anchors came out identical to Whisper's own line times.
Check the report before `--apply`: it warns when the last line lands past the
end of the song, which means the alignment slipped.
## When a song has NO lyrics, check LRCLIB again first
`lyrics-regenerate` only revisits songs that already have lyrics, and the
lyrics-worker transcribes anything with none — so a song can end up with a
Whisper transcript even though LRCLIB had the real words all along. That is
exactly what happened to "Take Me to the End". Before reaching for agy on a
freshly transcribed song, search LRCLIB by hand; if it has the words, the
`retime_lyrics.py` route above beats everything else.
Related: `lyrics-lookup` (LRCLIB first, then this), `lyrics-regenerate`
(replace wrong lyrics from LRCLIB), `deploy-prod`.