Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents
Abstract
Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this in-turn adaptation in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model conditional-continuation case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations conditioned on a teacher-forced, voice-converted prefix while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2% of cases, compared with 34.8% for PersonaPlex. An onset-limited reannotation that withholds Speaker A's response reproduces this comparison (69.6% versus 35.7% on 56 pairs). PersonaPlex otherwise continues unchanged (42.4%) or yields (22.7%). These conditional results illustrate the distinction between uptake and floor behavior; response-informed candidate selection, response scoring, and the conditioning setup limit attribution of the observed gap to PersonaPlex itself.