NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
The paper claims a single open speech model can listen, interrupt, call tools, and keep speaking in real time.
NemotronLabs VoiceChat uses a streaming speech encoder, decoder-only language model, specialized streams for agent text and function calls, an RNN-T branch for live transcription, and streaming TTS. The authors report strong full-duplex behavior, including 100% takeover after user interruptions and 93% response resumption after backchannels on their cited benchmarks. Tool use is less settled: the model reaches 82.5% tool-selection F1 on FDB 3.0, while argument accuracy and end-to-end execution still need work. Source: ArXiv · AI/CL/LG's note.
NemotronLabs VoiceChat uses a streaming speech encoder, decoder-only language model, specialized streams for agent text and function calls, an RNN-T branch for live transcription, and streaming TTS. The authors report strong full-duplex behavior, including 100% takeover after user interruptions and 93% response resumption after backchannels on their cited benchmarks. Tool use is less settled: the model reaches 82.5% tool-selection F1 on FDB 3.0, while argument accuracy and end-to-end execution still need work. Source: ArXiv · AI/CL/LG's note.
score 8