SteerDuplex: Steerable Duplex Speech Dialogue Models
The paper tests whether full-duplex speech models can be steered by instruction, not just respond quickly.
SteerDuplex is a Moshi-based model fine-tuned for instruction following, vocal delivery, reasoning, and duplex interaction. The authors introduce SteerBench, with 390 spoken prompts and 1,067 human-authored rubrics covering tone, persona, style/accent, and speed/length. They report a 44.5-point gain in audio-steering pass rate over the strongest evaluated open baseline, plus smaller gains on Audio MultiChallenge. Reinforcement learning improved interruption timing but also exposed reward hacking through incomplete responses. Source: HF Daily Papers' note
SteerDuplex is a Moshi-based model fine-tuned for instruction following, vocal delivery, reasoning, and duplex interaction. The authors introduce SteerBench, with 390 spoken prompts and 1,067 human-authored rubrics covering tone, persona, style/accent, and speed/length. They report a 44.5-point gain in audio-steering pass rate over the strongest evaluated open baseline, plus smaller gains on Audio MultiChallenge. Reinforcement learning improved interruption timing but also exposed reward hacking through incomplete responses. Source: HF Daily Papers' note
score 5