HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
HiPLEX splits real-time speech behavior into separate timing and content decisions.
The paper frames full-duplex dialogue as a control problem: when to speak, when to hold, and what to say. HiPLEX factors a pretrained policy into a control head for `pad`, `epad`, or `con`, and a content policy that selects tokens only when content is emitted. The authors route timing feedback and semantic feedback through separate advantage paths. In tests on Moshi seeds and Full-Duplex-Bench v1, they report fewer takeovers during pauses and backchannels, plus shorter post-interruption latency than GRPO, while keeping similar judged response quality. HF Daily Papers' note
The paper frames full-duplex dialogue as a control problem: when to speak, when to hold, and what to say. HiPLEX factors a pretrained policy into a control head for `pad`, `epad`, or `con`, and a content policy that selects tokens only when content is emitted. The authors route timing feedback and semantic feedback through separate advantage paths. In tests on Moshi seeds and Full-Duplex-Bench v1, they report fewer takeovers during pauses and backchannels, plus shorter post-interruption latency than GRPO, while keeping similar judged response quality. HF Daily Papers' note
score 5