$\tau$-Emotion-Bench: Exposing the Interaction Fragility of LLM Agents through Real-World User Emotions
Rishitosh K Singh ⋅ Seyyedamirhossein Saeidi ⋅ Shamanthak Hegde ⋅ Shrinidhi Kumbhar ⋅ Chitta Baral
Abstract
In long-horizon tool-calling tasks, LLM agents can fail not only because of tool-use errors, but also because of ineffective interaction with users. Yet existing benchmarks largely assume cooperative and behaviorally stable users, leaving the effect of emotional variation on agent reliability underexplored. We introduce $\tau$-Emotion-Bench, a benchmark for evaluating tool-calling agents under emotion-conditioned user interactions across four domains. To construct the benchmark at scale, we introduce Tracer, a scalable agentic data-generation pipeline that constructs policy-valid tool workflows, grounds them into executable tasks, incorporates user preferences and task-specific emotional context, and verifies task validity against the environment. The resulting tasks are mapped to a structured emotion space and paired with emotion-conditioned user simulators, enabling controlled evaluation while preserving the underlying task semantics. Experiments show that frontier models remain unreliable under these interactions, with substantial variation in task success and consistency across user conditions. Beyond evaluation, we use Tracer-generated data to fine-tune two models, resulting in consistent improvements on $\tau$-Emotion-Bench while maintaining comparable performance on $\tau$-Bench. Overall, $\tau$-Emotion-Bench provides a scalable framework for constructing and evaluating agentic tasks under diverse user behaviors.
Chat is not available.
Successful Page Load