Time correct lyrics from a machine transcript by aligning the two word streams
This commit is contained in:
@@ -114,5 +114,40 @@ Written into `data.tags` so a later run can tell where lyrics came from:
|
||||
| `from the web via agy (untimed) — check and Tap-sync` | agy, words only |
|
||||
| `from a file (<name>) (synced\|untimed)` | `--from-file` |
|
||||
|
||||
## Right words + real timings: `retime_lyrics.py`
|
||||
|
||||
The two sources fail in opposite ways — the web (LRCLIB plain, agy) has the
|
||||
right words and line breaks but no timings; Whisper has a time for every line
|
||||
but mishears words and breaks lines mid-phrase ("To show for the / years").
|
||||
Rather than re-rolling agy for a timed answer that may never come, align them:
|
||||
|
||||
```bash
|
||||
# correct words in a file (LRCLIB plain text, agy output, or pasted lyrics)
|
||||
python3 scripts/lyrics/retime_lyrics.py --id MU5dlCRTLY8 --words correct.txt
|
||||
python3 scripts/lyrics/retime_lyrics.py --id MU5dlCRTLY8 --words correct.txt --apply
|
||||
# take timings from a backup rather than what is live
|
||||
… --id ID --words w.txt --times-from ytplayer-lyrics-backup-….json
|
||||
```
|
||||
|
||||
It matches on a **word stream, not line to line**, precisely because the line
|
||||
breaks disagree: each transcript line's words are spread across the gap to the
|
||||
next line, `difflib` aligns the two word sequences, and a correct line takes the
|
||||
time of the earliest transcript word inside it. Times are forced monotonic, and
|
||||
lines that matched nothing are interpolated between their neighbours. On "Take
|
||||
Me to the End": 46 correct lines, 44 timed directly, 2 interpolated — and the
|
||||
anchors came out identical to Whisper's own line times.
|
||||
|
||||
Check the report before `--apply`: it warns when the last line lands past the
|
||||
end of the song, which means the alignment slipped.
|
||||
|
||||
## When a song has NO lyrics, check LRCLIB again first
|
||||
|
||||
`lyrics-regenerate` only revisits songs that already have lyrics, and the
|
||||
lyrics-worker transcribes anything with none — so a song can end up with a
|
||||
Whisper transcript even though LRCLIB had the real words all along. That is
|
||||
exactly what happened to "Take Me to the End". Before reaching for agy on a
|
||||
freshly transcribed song, search LRCLIB by hand; if it has the words, the
|
||||
`retime_lyrics.py` route above beats everything else.
|
||||
|
||||
Related: `lyrics-lookup` (LRCLIB first, then this), `lyrics-regenerate`
|
||||
(replace wrong lyrics from LRCLIB), `deploy-prod`.
|
||||
|
||||
Reference in New Issue
Block a user