Belief Attribution Requires Behavioral Evidence and Agents Now Supply It
Abstract
Machine learning attributes beliefs to language models through verbal reports, output probabilities, probing classifiers, and localized circuits, and these operationalizations disagree. Philosophical work on belief in language models has explained why measurement moved inside the model: a model that only completes text offers little behavior to interpret, while its activations are fully accessible. We argue that this premise is contingent and that agentic deployment has removed it. A model that acts through tools over many steps, under consequences it can infer, faces choices under uncertainty with stakes, which is the setting in which Ramsey's functional account of belief was written. We give a criterion for belief attribution indexed to a decision class, a perturbation family, and a stake structure, built on recent utility-agnostic conditions for eliciting beliefs from an agent's choices. Structural evidence such as a probe direction counts only when intervening on the structure moves behavior across the perturbation family. Four published results, on contrast-consistent search, causal-tracing localization, theory-of-mind tests, and alignment faking, are reread as verdicts returned by single clauses of the criterion. We close with a protocol for measuring what an agent believes.