Uncovering Degradation in Full-Duplex Speech Models via Long, Topic-Driven User Simulation
Abstract
Full-duplex speech-to-speech systems are increasingly capable of human-like conversational behavior, but whether that behavior holds up over long conversations remains untested. Existing evaluations either use a fixed multi-turn audio and score only the final turn, or test the system with a prompted user-simulator whose own consistency across turns is not well controlled. We present a framework for measuring multi-turn degradation in full-duplex systems, built on a topic-driven user simulator grounded in a long reference dialogue chain, which gives explicit control over persona, topic progression and responsiveness. Since naturalistic corpora with such long multi-turn dialogues are scarce, we construct these chains from short dyadic dialogues by scoring speaker initiative, clustering speaker roles, and chaining topically related dialogues. The simulator drives a live conversation with the target system using a turn-taking model, reference-steered user-LLM, and expressive speech synthesis. We score every system turn along three axes grounded in the dialogue literature across turn-taking, competence and engagement taxonomy. We further propose turn-anchor rules to determine multi-turn evaluation segments for full duplex audio. We run 100 simulations per system, up to 10 minutes each, across three open-source full-duplex speech models. All three models show consistent degradation over the course of a conversation on all three axes, and we detect points of complete model breakdown. Our experiments indicate that PersonaPlex is the most robust, with all engines collapsing sharply late in the conversation as they increasingly stop responding.