Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
A lightweight ASR head lets a duplex speech model transcribe the user in real time without rebuilding the model.
The paper adds streaming transcription alongside the model’s existing agent text output. The authors say it preserves full-duplex behavior, including turn-taking and barge-in handling. In tests inside the duplex setup, the method reports 10.21% streaming average WER on the HuggingFace Open ASR Leaderboard. Used as a standalone streaming ASR model, the same architecture reaches 7.73% WER, and the authors say they will open-source training and inference code. ArXiv · AI/CL/LG's note
The paper adds streaming transcription alongside the model’s existing agent text output. The authors say it preserves full-duplex behavior, including turn-taking and barge-in handling. In tests inside the duplex setup, the method reports 10.21% streaming average WER on the HuggingFace Open ASR Leaderboard. Used as a standalone streaming ASR model, the same architecture reaches 7.73% WER, and the authors say they will open-source training and inference code. ArXiv · AI/CL/LG's note
score 5