Intent-Preserving Adversarial User Simulation for Multi-Turn Agent Evaluation
Atri V Sharma ⋅ Brian Formento ⋅ Alessio Lomuscio
Abstract
Multi-turn agents are commonly evaluated with LLM-based simulated users, which can be overly cooperative and can therefore overestimate an agent's performance. We thus introduce an intent-preserving adversarial user simulator, a framework for stress-testing agents with difficult but valid user interactions. Our method optimizes the simulator's high-level adversarial plan for each task instance. Candidate planner prompts are evolved with a genetic prompt-optimization loop that uses rollout traces, granular task-completion signals, and validity feedback to favor conversations that reduce task success without changing the user's underlying goal. We evaluate our approach on the MultiWOZ dataset and the retail and airline domains of $\tau$-bench with two target LLM agents, comparing against standard simulators and hand-designed non-collaborative user behaviors, and demonstrate the effectiveness of our approach in exposing robustness gaps in both models (decreasing pass^4 by up to $69\%$ on MultiWOZ and 100\% on $\tau$-bench), while maintaining goal preservation performance. A human validation study shows that our automated validity judgments are well calibrated with human annotators, and an analysis of the optimized planner prompts reveals recurring adversarial strategy families, such as conditional and staged information disclosure. Our results show that validity-aware adversarial simulation can expose robustness gaps missed by conventional user simulators.
Chat is not available.
Successful Page Load