Runtime detection of faulty reasoning in 2B models without a large judge
Abstract
Deploying a 2B model is a decision about inference cost, yet the uncertainty methods used to decide whether its outputs can be trusted are the ones that cost the most. Sampling-based detection multiplies the inference bill by the sample count, and log-probability methods need output probabilities that frontier APIs increasingly withhold. A locally deployed small model is the one regime where the per-token distribution is already computed and free, while the standard alternative, escalating suspect outputs to a large model, is exactly the cost such a deployment set out to avoid. We ask whether the free signal can stand in for the large judge. On two 2B open-weight models answering GPQA-Diamond, a 50-token windowed entropy statistic with a per-trace threshold matches an 80B-A3B judge at discriminating faulty from sound reasoning on the same flagged spans (Youden's J 20.6 against 15.3 at an untuned median cut, within noise at this n), at effectively zero marginal cost and with no labelled calibration set, and the same direction holds on whole-trace entropy for all nine models of a wider bank. The detector has a measured ceiling, 85 of 350 traces are calm and wrong, and the judge is unreliable as a critic. It calls 83.7% of correct reasoning faulty, critique-and-retry yields no reliable gain over a content-free nudge, and a model handed a complete answer adopts it on 91% of items, errors included. The comparison rests on an evaluation that labels every response with an auditable outcome category and its cost. Under it, on the 86 physics items, the two subjects, 3.5 accuracy points apart, differ in failure composition (p = 0.0037), and chain-of-thought's twenty points cost 350-550x the tokens. At 2B the free signal suffices for detection, and the large judge is neither needed for it nor able to repair.