Megadose AI progress, ranked and analyzed.

Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech

· ArXiv · AI/CL/LG ·
WER made the systems look closer than they were; switch-point metrics exposed the gap.

The paper evaluates eleven ASR and audio language models on a 2,000-utterance English-Yoruba code-switched set. Its central result is that the best WER system and a leading audio LM are statistically close on WER, while the audio LM is better on every switch-localized measure. Yoruba recognition largely breaks down across faithful systems, with errors clustering at switches into Yoruba. Some generative audio LMs also fail as exact transcribers, adding translation, verbosity, or prompt leakage depending on the prompt. ArXiv · AI/CL/LG's note

score 4

Categories: Research