arXiv · 2610.08683
Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue
Abstract
Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner's turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lichen Zhu, Yueqian Lin, Yiheng Wang, Hai "Helen" Li, Yiran Chen. 2026-10-06. Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue. https://arxiv.org/abs/2610.08683
Cite the original work for its findings. Save a collection to share your selection of sources.