Megadose AI progress, ranked and analyzed.

SteerDuplex: Steerable Duplex Speech Dialogue Models

· HF Daily Papers ·
The paper tests whether full-duplex speech models can be steered by instruction, not just respond quickly.

SteerDuplex is a Moshi-based model fine-tuned for instruction following, vocal delivery, reasoning, and duplex interaction. The authors introduce SteerBench, with 390 spoken prompts and 1,067 human-authored rubrics covering tone, persona, style/accent, and speed/length. They report a 44.5-point gain in audio-steering pass rate over the strongest evaluated open baseline, plus smaller gains on Audio MultiChallenge. Reinforcement learning improved interruption timing but also exposed reward hacking through incomplete responses. Source: HF Daily Papers' note

score 5

Categories: Research