Towards Quantifying Benchmark Optimization in ASR Models
Top open-source ASR models can reproduce benchmark transcript spans even when the audio does not support them.
The paper proposes probes for benchmark optimization where speech audio is contradictory, masked, or ambiguous. In those cases, the strongest models still output verbatim benchmark reference text. The authors say mechanistic tests suggest narrow acoustic cues can push models toward a benchmark-conditioned transcription policy. They also report that this behavior can be manipulated with low-rank steering or, in some cases, by appending audio to a segment. HF Daily Papers' note
The paper proposes probes for benchmark optimization where speech audio is contradictory, masked, or ambiguous. In those cases, the strongest models still output verbatim benchmark reference text. The authors say mechanistic tests suggest narrow acoustic cues can push models toward a benchmark-conditioned transcription policy. They also report that this behavior can be manipulated with low-rank steering or, in some cases, by appending audio to a segment. HF Daily Papers' note
score 5