When Speech Errors Break Tasks: Cross-Layer Evaluation of Conversational Voice Agents
Abstract
We introduce a cross-layer protocol for interpreting voice-agent behavior at the speech-task interface. Pragmatic Intent Failure (IntentFail) records whether a version-pinned router assigns different routes to a reference transcript and its ASR hypothesis. Slot & Entity Error Rate (SEER) tests whether typed reference values remain recoverable after spoken-form normalization and separates mutation from drop. Dialogue State Failure and entity lifecycle measures extend the same trace representation across composed trajectories. The evaluation suite is grounded in recurring industry structure through a governed process that distills handler relations, intent cues, dialogue patterns, and entity formats from access-controlled real-world interactions, instantiates new records with fictional values, and applies task-specific quality gates. Across six ASR systems on 336 speech-rendered turns, lexical and routing diagnostics yield different nominal leaders, although routing confidence intervals overlap. Pooled over 3,996 scored pairs, 49% of fixed-router changes occur at zero normalized WER. On 400 entity-rich turns, tracking-number SEER ranges from 18.8% to 34.4% and order-ID SEER from 7.2% to 44.6%, compared with 0.4% to 3.0% for dates. Manual analysis of 60 target-gated route changes identifies recurring critical-word substitutions and sensitivities to orthographic and punctuation form. Experiments with 462 composed dialogues distinguish turn-local from terminal-state behavior, while a synthesis-boundary probe extends the trace across generated speech. Together, these measurements turn aggregate errors into localized and countable behavioral patterns.