Transcribe a sung vocal into lyrics with word-level timings (Whisper large-v3 on GPU). Use a vocal from a previous stems test or upload one. This is speech-to-text only — phoneme alignment is a later pipeline stage.
start_ms/end_ms. Best results on a clean vocal stem (use the stems service first). With lyrics, the corrected words inherit the ASR word times (transferred via a word-level delta, not re-aligned to audio) and become the baseline when they fit; ASR still runs to verify (a delta + low-confidence list is returned). Singing ASR is imperfect — the editor allows correction downstream.