AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks
Abstract
Speech-to-speech (S2S) language models are increasingly deployed as voice agents that schedule appointments, take grocery orders, and plan events. Reliable deployment in such settings requires accurate listening, grounded tool use, state tracking, knowledge-base access, and recovery from user corrections across many turns. Existing voice agent benchmarks often score final database state or per-call tool accuracy, which can obscure the first failure point when an early mistake cascades through the rest of a conversation. We introduce AudioAgentBench, a fixed-trace benchmark suite with six multi-turn voice-agent tasks, 221 turns of pre-recorded audio, golden tool calls, tool schemas, knowledge bases, and matched TTS and human-recorded variants. The fixed audio traces make turn-level expectations reproducible and support a diagnostic scoring protocol that attributes failures across tool use, state tracking, ambiguity handling, knowledge grounding, and instruction following. An LLM judge applies the protocol to each turn by comparing model outputs against golden tool calls, task state, and task-specific knowledge bases, while using cross-turn realignment, conditional penalty absorption, and category-aware dimension gating to avoid double-counting cascading errors. We evaluate seven S2S models across 420 continuous session runs. The top model averages an 80.3\% pass rate, but the top four models have overlapping 95\% confidence intervals, and no model leads on more than two of six tasks. On a 408-run turn-level diagnostic subset with 14,907 judged turns, tool-use accuracy on turns where a tool call was expected peaks at only 61.4\% and drops to 32.3\% on a confusable-name scheduling task. Errors cascade by up to 3.15× in stateful tasks, and tool-use accuracy on a 30-turn grocery task falls by 40.5 percentage points from the first half to the last half. Scoring only turns where a tool call was expected reduces apparent tool-use accuracy by 22–44 percentage points across models, which reveals a larger gap between strong and weak systems than aggregate task success suggests. AudioAgentBench surfaces action, state, and recovery failures that final success metrics obscure, providing a finer-grained diagnostic benchmark for production-oriented voice agents.