Megadose Built for builders and researchers.

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

· ArXiv · AI/CL/LG ·
Full-duplex speech models coordinate their turns, but they hand off the floor much later than humans.

The paper tests two PersonaPlex-7B instances exchanging audio tokens in unscripted dialogue on a shared clock. Their timing is coupled, since re-pairing speakers across conversations breaks the pattern. But median floor transfer lands at 400-560 ms, versus 137 ms in Switchboard human speech. The authors say delayed channels shift responses one-for-one, suggesting the models wait reactively after a perceived ending rather than projecting the turn end.

Source: ArXiv · AI/CL/LG's note

score 4

Categories: Research