Megadose AI progress, ranked and analyzed.

StepAudio 3 Realtime Technical Report

· HF Daily Papers ·
StepAudio 3 Realtime is pitched as a speech model that can reason while it is already talking.

The paper describes a continuous listen-converse-think-act loop for real-time spoken interaction. Its duplex system is meant to handle pauses, backchannels, and interruptions without breaking the exchange. The authors say Think-While-Speaking lets private reasoning run in parallel with spoken output, preserving low latency. Reported results include 73.0 on StepAudioChat in reasoning mode, 90.6 on MMSU, 98.9 on Artificial Analysis Full-Duplex Bench, and 56.0% macro task success on τ-Voice. HF Daily Papers' note

score 7

Categories: Model Releases, Research