Megadose AI progress, ranked and analyzed.

SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models

· ArXiv · AI/CL/LG ·
Sharding the same math problem across spoken turns made commercial speech systems less accurate.

SCB tests 103 GSM8K problems as full, concatenated, and incremental spoken disclosures. The paper reports sharded accuracy drops of 5.0 to 25.3 points versus concat across four commercial systems. LEGO, an internal SCBX speech pipeline with explicit context management, held 77.5% accuracy in all three settings. GPT-4o Realtime is listed at 76.6% in the sharded condition. ArXiv · AI/CL/LG's note

score 5

Categories: Research