Reliable User Simulation in Complex Multi-Step Workflows
Abstract
LLM-based user simulators are increasingly used to evaluate conversational assistants, but simulators can drift from their assigned goals, causing evaluations to attribute simulator errors to the assistant. To study simulator fidelity, we introduce a 500-workflow evaluation set formalized as hierarchical task networks. The workflows combine sequential tasks, cross-task dependencies, and fallback options, requiring simulators to track their progress across turns. Under zero-shot prompting, fidelity varies substantially across model tiers and declines as workflow complexity increases. Most failures arise from goal drift, such as skipping options or switching prematurely to a fallback. We further show that assistant response style is a confounder: the same simulator receives substantially different fidelity scores across assistant response styles. We propose ReUST (Reasoning + User State Tracking), which carries dialogue state across turns. ReUST improves Goal Success Rate by more than 12 percentage points on Medium and Large models and substantially reduces sensitivity to assistant response style.