Localizing and Causally Validating Argument-Level Hallucination in LLM Tool Calls
Abstract
Large language model agents that call external tools sometimes populate arguments with values unsupported by the conversation context, a failure we term argument-level hallucination. Prior detection methods judge an entire tool call at once and rely on correlational evidence rather than showing the detected signal actually drives model behavior. We localize hallucination to individual arguments using internal activations at each argument's token position, trained via a causal-contrastive procedure that contrasts matched contexts with and without supporting evidence. A linear probe over these activations separates grounded from ungrounded arguments with near-ceiling accuracy, and we validate the mechanism causally: suppressing the identified direction recovers grounded behavior rather than merely correlating with it, an effect specific to the true grounding direction, as shown by random-direction and label-shuffled controls. Across seven models spanning multiple scales and architectures, the correlational signal stays consistently strong, while causal actionability varies in a structured way: strong from a single-layer intervention at small scale, recoverable only via multi-layer intervention at larger scale within a family, with the required layer count itself increasing with model size, and, in one model, negligible under any intervention we test -- an open exception, not explained by data volume, layer count, or architecture family, rather than a resolved boundary of our claim.