Does the Voice Matter? Frustration Detection on Real Clinical Voice-Agent Calls
Abstract
Autonomous voice agents now conduct routine clinical follow-up calls, and because no service can review every call by hand, some automated triage is unavoidable. A patient becoming frustrated is one thing that triage should catch. These systems typically read only the transcript the speech-to-text stage produces, not the voice, and whether that is enough is largely untested on real clinical speech. We study per-turn frustration detection on real patient calls from a deployed clinical voice agent, with turn-level labels agreed by three annotators and a length-matched acted corpus as a control. Patient turns are overwhelmingly short: across the deployed calls, roughly 59% are one or two words, against 13% in the full acted corpus, and a corpus accuracy gap remains even after the acted set is length-matched. On one-word turns, no transcript-only condition excludes chance: no interval clears 50%, recall on frustrated one-word clinical turns is 9.5% for all three models, and a classifier given only the word count reaches 58.0%. Audio does not fail in the same way, but its value depends on the model. Across two frontier models, audio improves clinical-call accuracy by about 4.0 percentage points, reaching significance (4.9 percentage points) only when the acted control is pooled in, while a smaller open model performs 12.7 percentage points worse with audio than with the transcript. The failure is also silent: of the turns confidently cleared using transcripts, 37.9% were in fact frustrated, compared with 17.8% using audio in our balanced sample. Hearing the voice therefore roughly halves the rate of silently cleared frustration, although the absolute rates reflect the constructed base rate rather than deployment prevalence. A short patient turn should not be cleared on the transcript alone.