When Do Facts Become Evidence? Evidentiary Threshold Signatures for Interpreting Agent Behavior
Abstract
Facts do not become evidence at the same point for every LLM agent. When LLMs are used for judgment tasks, understanding how they use evidence matters as much as whether they reach correct conclusions. We introduce a behavioral framework for auditing evidence use in LLM judgment. The framework decomposes evidence use into five observable components (fact identification, relational bridging, sufficiency judgment, label selection, and commitment strength) and localizes model disagreement through controlled perturbations. Experiments on 10,150 task records from DeepSeek and Doubao/Seed, with controlled assays across all three models including Qwen, reveal that: (1) relational-chain disruption yields the largest and most consistent cross-model effect, despite near-constant fact coverage; (2) all models share an assay-defined change point when target relations become explicit, yet exhibit distinct pre-transition policies: some infer support from implicit relational cues while others withhold commitment; (3) exploratory activation patching in Qwen3-4B uncovers a causal signature concentrated in middle-to-late layers. We formalize these response profiles as Evidentiary Threshold Signatures (ETS), whose compact parametric summary captures threshold location and transition sharpness. ETS offers a principled tool for auditing and comparing evidence-acceptance policies under controlled conditions.