Megadose AI progress, ranked and analyzed.

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

· ArXiv · AI/CL/LG ·
MTVA-Bench tests the language model in a cascaded voice agent under phone-call conditions, not as a standalone chatbot.

The paper introduces a benchmark with 49 agents, 490 reviewed scenarios, and support for 7 languages. It uses an LLM-simulated caller, a mock backend, deterministic tool-call checks, and two LLM judges that must cite transcript evidence. In a seven-model study, correct tool choice was relatively close across six models, but overall scores differed by 24.4 points because of arguments, action order, rule compliance, and surrounding dialogue. ArXiv · AI/CL/LG's note

score 6

Categories: Research