When Agents Trust the Wrong Tool Output: Acceptance Across Corruption Types and a URL-Host Ablation
Abstract
Tool-using language-model agents must decide whether to use an output after a tool call returns. We isolate this post-tool decision with four actions: accept, verify, abstain, and escalate. On a paired web benchmark grounded in 50 public sources, baseline Gemini 2.5 Flash and GPT-5.4 mini each directly accept 11/19 same-entity false-value corruptions but 0/15 wrong-entity substitutions; Gemini 3.1 Pro accepts 3/19 and 0/15, respectively. We then change only the visible URL host in a pre-specified ablation. For GPT-5.4 mini under two caution-oriented prompts, replacing real domains with example.org eliminates 4/19 and 5/19 same-entity accepts, but raises deferral on reliable outputs from 1/50 to 33/50 and 30/50. No baseline or Gemini Flash condition shows the same clear effect. Thus, within this benchmark, acceptance differs across type-defined subsets, sensitivity to a conspicuous provenance cue is model-policy-specific, and a lower unsafe-accept rate can conceal a much larger immediate-use cost.