Diagnosing Agent Failure: A Behavioral Taxonomy for LLM User Simulators
Abstract
As autonomous agents are increasingly deployed, the community relies on simulated users to evaluate them. However, current simulator metrics focus heavily on surface-level text properties rather than accounting for how the agent actually executes on the simulated user's input, which is ultimately what matters most. To bridge this gap, we present a comprehensive 25-metric behavioral taxonomy (capturing constructs like frustration, escalation, and information exchange) designed to measure how users interact with agents. We demonstrate that our taxonomy captures human conversational structure far better than prior baselines (averaging over 76.1 +/- 2.2 Sørensen--Dice alignment across domains versus the baseline's 51.8 +/- 2.4) and provides stronger, more consistent predictions of synthetic task failure (yielding significant outcome associations for up to 52% of our metrics compared to a maximum of 25% for the baseline, while delivering equal or higher joint predictive power). Crucially, however, an interaction analysis across human and synthetic conversations exposes a behavioral divergence: even when measured with strong metrics, LLM simulators experience failure differently than real users. Simulators lack the dynamic adaptability of humans, exposing a gap where real interactions are far more complex to model. We conclude that we cannot assume simulated failures perfectly mirror real-world failures. Nonetheless, our metrics offer both a stronger standard for predicting synthetic failure and a precise diagnostic vocabulary for measuring where current simulators still fall short.