DUAL-AudioBench: Closed-Loop Evaluation of Belief-State Synchronization in Audio-Conditioned Voice Agents
Abstract
A voice agent files a booking change, the call pauses, and the caller returns ten minutes later. Nothing said aloud reveals whether the change went through. The agent has to derive it from the action it took, the time that passed, and what the caller now reports. Existing audio benchmarks score the endpoint, where a stale belief paired with a lucky action is indistinguishable from a correct inference. We introduce DUAL-AudioBench, a novel turn-based audio-replay benchmark of 84 scenarios in 42 matched pairs across 14 domains. The two branches of a pair, aligned and misaligned, differ in one early fact that changes the hidden outcome while the post-gap words and the action menus stay fixed. We score both the agent's state probabilities and chosen action. Across 4,368 trajectories with three audio-capable models the failure is asymmetric. On the misaligned clue branch, where the intended process fails under the correct pre-gap action, final-action accuracy is 39.3 to 60.7%. On the aligned clue branch, where it succeeds under that action, and the correct move is simply to close the request, accuracy collapses to 4.8 to 33.3%. These agents notice trouble and cannot conclude that the absence of trouble means the work is done. Two of the three show essentially no positive belief revision toward the state that actually holds after hearing the resumed evidence, and the third moves 0.309. Supplying the realized state in plain language lifts aligned-branch accuracy to 44.0 to 78.6%, showing that mistaken state inference explains much of the failure. However, even with the state provided, two models remain no better than a constant-action baseline, indicating an additional difficulty in mapping state to action. We therefore compare performance against baselines computed from each model's realized trajectories.