Detecting and Correcting Reasoning Failures: A Controlled Negative Result for a Residual-Stream Monitor
Abstract
Language models often solve math problems with fluent reasoning that still ends in a wrong answer. If the model’s internal state signaled the error while the solution was being written, an agent could stop and recompute before committing to it. We test the most direct version of this idea: track how far the hidden state has drifted from where the solution started and, when it drifts too far, roll back a few tokens and prompt the model to recheck its work. We evaluate the monitor by whether the final answer is correct, in the model’s intended prompt format, and against controls that isolate what it adds, including a no-intervention baseline and intervening at a random point. Across the evaluated configurations on two open 7B models and two arithmetic benchmarks, the monitor provides no consistent accuracy improvement. In the two configurations with the most incorrect answers, detection AUC is not distinguishable from chance; false-trigger rates reach 98%, and monitor-guided intervention shows no recovery advantage over random-position intervention. We also show why letting the model judge its own coherence makes such a method look successful when it is not, and we state the conditions a mid-generation monitor must meet before it can be trusted.