Knowing Before Saying: Internal Correctness Signals in Mathematical Reasoning
Mei Chen ⋅ Gordon Stein ⋅ Andreas Ziegler
Abstract
Language models are often confidently wrong. We ask whether a model's activations carry a better signal of its own errors than its stated confidence does. We test four open models on four tasks: MMLU-Pro, MATH, GPQA-Diamond, and a geometry task graded by a compiler. We put a linear probe on the hidden state to predict whether an answer is right. This probe was better than the model's stated confidence in 12 of 16 cases, by 0.09 AUROC on average. The gap is largest on MATH. The probe is not just detecting hard questions: it still works when the same question is attempted five times. On Mistral and MMLU-Pro, the signal exists before the model is asked how confident it is. On Mistral, the model uses this signal when it reports confidence. When we erase the signal from the hidden state, the stated confidence no longer separates right answers from wrong ($0.84 \to 0.56$), while erasing a random direction did nothing. On GLM, erasing it changes nothing. On Qwen3.6 the stated confidence carries almost nothing to erase, and amplifying the signal adds a little. In short, four open models carry more confidence signal about their own errors than they report. In Mistral, we have shown this direction is used when it self-reports confidence.
Chat is not available.
Successful Page Load