MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents
MSI-Bench finds current voice agents still break down when several people are speaking and instructions are speaker-scoped.
The benchmark uses 1,152 short multi-party audio scenes in English and Mandarin, with expected tool calls and atomic rubrics. The best reported systems clear all rubrics on 66.8% of English cases and 54.5% of Mandarin cases; the strongest open-weight setups fall to 34.0% and 19.3%. The paper separates audio perception failures from reasoning failures, saying open-weight models are held back by the multi-speaker front end while frontier systems still miss speaker-scoped decisions even on clean transcripts. It also notes a recurring restraint problem: models often answer when no one has addressed them. ArXiv · AI/CL/LG's note
The benchmark uses 1,152 short multi-party audio scenes in English and Mandarin, with expected tool calls and atomic rubrics. The best reported systems clear all rubrics on 66.8% of English cases and 54.5% of Mandarin cases; the strongest open-weight setups fall to 34.0% and 19.3%. The paper separates audio perception failures from reasoning failures, saying open-weight models are held back by the multi-speaker front end while frontier systems still miss speaker-scoped decisions even on clean transcripts. It also notes a recurring restraint problem: models often answer when no one has addressed them. ArXiv · AI/CL/LG's note
score 6