SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models
Sharding the same math problem across spoken turns made commercial speech systems less accurate.
SCB tests 103 GSM8K problems as full, concatenated, and incremental spoken disclosures. The paper reports sharded accuracy drops of 5.0 to 25.3 points versus concat across four commercial systems. LEGO, an internal SCBX speech pipeline with explicit context management, held 77.5% accuracy in all three settings. GPT-4o Realtime is listed at 76.6% in the sharded condition. ArXiv · AI/CL/LG's note
SCB tests 103 GSM8K problems as full, concatenated, and incremental spoken disclosures. The paper reports sharded accuracy drops of 5.0 to 25.3 points versus concat across four commercial systems. LEGO, an internal SCBX speech pipeline with explicit context management, held 77.5% accuracy in all three settings. GPT-4o Realtime is listed at 76.6% in the sharded condition. ArXiv · AI/CL/LG's note
score 5