From Text to Voice: How Real-Time Pipelines Change Interruption Recovery in Cascaded Clinical Agents
Abstract
Clinical voice agents are commonly built as a cascade of speech-to-text, a language model, and text-to-speech, and real patients interrupt them. Once a turn is cut short the recovery is a text problem, but a live pipeline decides what truncated transcript the model then sees: it keeps playing for a moment after the patient starts, so the barge-in policy governs how much of the unfinished turn is written to the agent's record. We plan a clinical interruption once and replay it identically in text and through a live voice pipeline under three barge-in policies. We score from the patient's side whether the clinically required content survives. On the same material, text and voice verdicts sometimes coincide, both a consistent pass or both a consistent failure, and sometimes diverge; which happens turns non-monotonically on the model, the wording of the interruption, and the barge-in policy. As little as one or two words the pipeline commits after the patient starts speaking can flip the safety verdict in either direction, and a defensive prompt that tells the agent to finish its answer does not prevent this. We also test a simple change: making the agent aware of the interruption. For the recovery turn we add to what the model sees a short description of what happened, what it had said, what was spoken over the patient, and what it was cut off before saying. It improved recovery under some interruptions and worsened it under others, so such an intervention must be measured before it is trusted.