What Did I Just Say? Self-Listening for Full-Duplex Speech Models
The paper targets a mismatch in full-duplex speech models: the model may not know what audio the user actually heard before an interruption.
The authors call this “anchor interruption,” where recovery depends on the last realized speech, not just generated text. Their Self-Listening approach feeds the model’s played speech back in alongside user speech and model text. They also introduce AnchorSpeech to test whether models stay consistent with the last completed spoken item. Experiments reported in the paper show better anchoring than full-duplex baselines. HF Daily Papers' note
The authors call this “anchor interruption,” where recovery depends on the last realized speech, not just generated text. Their Self-Listening approach feeds the model’s played speech back in alongside user speech and model text. They also introduce AnchorSpeech to test whether models stay consistent with the last completed spoken item. Experiments reported in the paper show better anchoring than full-duplex baselines. HF Daily Papers' note
score 6