Stress-Testing Hidden-State Error Probes in LLMs
Abstract
Recent work has shown that linear probes can decode reasoning errors from the hidden states of language models, motivating the interpretation that models possess a latent form of error awareness. However, high decodability alone does not establish what a probe is reading, whether the signal generalizes beyond a particular task, or whether activation interventions are causally specific. We stress-test this interpretation using matched clean/error reasoning pairs derived from two public benchmarks: GSM8K for arithmetic reasoning and ProofWriter for deductive reasoning. Across Qwen2.5-1.5B-Instruct and SmolLM2-1.7B-Instruct, we measure error separability at the input, after self-attention, and after the MLP of the first Transformer block. On 160 GSM8K pairs, AUROC rises from 0.500 at the block input to 0.698/0.687 after attention and 0.704/0.678 after the MLP. In contrast, the same analysis remains near chance on 120 ProofWriter pairs, and cross-domain transfer is also near chance. Moreover, on the evaluated baseline subset, step surprisal reaches 0.948 and 0.910 AUROC, substantially higher than the corresponding full-dataset Block-0 probe AUROCs. Because these estimates use different sample sizes, we treat surprisal as a strong alternative-explanation diagnostic rather than a controlled head-to-head comparison. Finally, matched clean activation patches produce almost the same downstream probe-score changes as mismatched clean patches from unrelated problems. We argue that decodability and intervention magnitude alone are insufficient evidence for a domain-general, causally specific representation of reasoning-error awareness. We propose three falsification criteria for future claims: cross-domain generality, resistance to simpler alternative predictors, and causal intervention specificity.