Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
The paper says ASR errors can be driven less by recognition than by uncontrolled transcript style.
The authors argue modern ASR systems mix verbatim and intended transcription policies, which can distort WER and word timing. Their task-token method lets a model choose the style, lifting German disfluency F1 from 10% to 79% zero-shot despite English-only training. They also report improved word-level timestamps on disfluent speech using supervised cross-attention fine-tuning. HF Daily Papers' note
The authors argue modern ASR systems mix verbatim and intended transcription policies, which can distort WER and word timing. Their task-token method lets a model choose the style, lifting German disfluency F1 from 10% to 79% zero-shot despite English-only training. They also report improved word-level timestamps on disfluent speech using supervised cross-attention fine-tuning. HF Daily Papers' note
score 5