Effectiveness of Hidden-State Probes for Localizing Reasoning Failures
Abstract
When the final answer from a long reasoning trace is verified as wrong, the verification does not show whether the problem lies beyond the model's capability, whether another attempt would succeed, or which step in an otherwise viable trace made it fail. Standard hidden-state probes are trained on final-answer correctness, which does not distinguish these cases. Monte Carlo rollouts distinguish the cases and locate the failing step at additional inference cost. We test how much of this the hidden states of 4B and 8B reasoning models carry and compare each probe with the strongest baseline that uses no hidden states. At the end of an incorrect 8B trace, a calibrated correctness probe at a midpoint threshold misses 82\% of cases in which a single step makes the trace unrecoverable and 52\% of cases in which another attempt would succeed. A probe with a different training design over the same states reverses this ordering, so either ordering can be read from those states. At both scales, no tested probes locates the failing step more reliably than a rule based only on trace length, and at 4B the best probe fails to outperform this rule in any of 10 paired runs on identical traces. A calibrated correctness probe can therefore inform a retry decision, but the failure-type ordering it reports still depends on how it was trained.