Causal Probes for Interpreting Agent Behavior
Abstract
The only record of what an agent did is a trace it wrote about itself, and where that trace pairs reasoning with a matching action, the pairing is read as an account of the behavior. That account does not survive intervention. Checkpoint resumption rebuilds the model's context at a stored reasoning block, substitutes the reasoning, and samples only the next action, so a counterfactual at one step of a long-horizon run costs a single completion rather than a full re-execution. Across 324 trajectories from Agents' Last Exam, corrupting a non-recoverable derived value moves the next command no further than a meaning-preserving paraphrase, at −2.9 points on Opus (95% CI [−14.5, +8.8]), while replacing a stated intention changes the action in 90% of Opus and 80% of GPT cases against identity floors of 20.0% and 22.9%. A crossed intervention adds a second dissociation: the same requested behavior draws 87.5% obedience written as the model's own reasoning and 20.0% as an external directive. Observationally the models look alike, reasoning naming the subsequent action in 25.6% to 31.8% of cases, yet content equally explanatory in the trace can arise from mechanisms differing by an order of magnitude, so inspection alone cannot determine which parts of a reasoning chain steer behavior.