Same Failure, Different Cause: Ungrounded Tool-Call Provenance in Simulated Agent Evaluation
ABHIMANYU PRASAD
Abstract
Tool-calling agents evaluated against LLM-simulated users frequently invoke tools with identifiers that no prior tool result ever confirmed. We show that this single surface failure has different underlying causes in different models, and that one of those causes would be concealed by a more faithful user simulator. On three matched tasks in $\tau^2$-bench retail, with both agents facing the same user simulator, Qwen2.5-7B-Instruct makes ungrounded tool calls at a conditional rate of 93.3%, and 79.6% of its ungrounded identifiers are copied verbatim from something the simulated user said. Llama-3.1-8B-Instruct fails less often (58.4%) but 74.5% of its ungrounded identifiers are fabricated with no traceable source. An agent that acts on a user-stated identifier without verifying it is only wrong when the user is wrong. Our simulator was wrong often: 86.5% of calls carrying a user-stated identifier returned an error, which is why the failure is visible here at all. A simulator drawing identifiers from the task database would make every such call succeed, and an agent that never verifies would be indistinguishable from one that always does. We quantify the exposure this depends on: the fraction of a simulator's stated identifiers that reach a tool call ranges from 46-89% for a 7B user model down to 0-21% for a 1.5B one on the same tasks. Simulator choice therefore determines how much material exists for this failure mode to appear at all. The direction holds across all three matched tasks in a two-model characterization over seven conversations.
Chat is not available.
Successful Page Load