Confidence Is Not a Stopping Rule: What a Self-Refinement Harness Can Observe About Its Own Progress
Abstract
A self-refinement loop is a minimal agentic harness: one model composed with itself in two roles, solver and critic, wrapped in control flow that decides after each revision whether to iterate again. Every such harness embeds a stopping rule, and the observable most readily available to it is the agent's own stated confidence. We ask a harness-engineering question: which observables of a running loop actually support the stopping decision? Using a transition-level protocol over 446 math and multi-hop QA questions and 1,784 stopping decisions with an open-weight model, we find that introspective observables are unusable while process-level ones are not. Stated confidence is saturated (98% of reports >= 95 against 41-50% accuracy), and conditioned on the current answer being wrong it does not rank the fixable cases above the unfixable ones (AUC = 0.485, 95% CI [0.466,0.506]), which we report as an equivalence-style bound: any monotone use of it has AUC-equivalent at most 0.534. The critic role is not uninformative (its verdict predicts harmful flips at AUC = 0.665), but a stopping rule reading that verdict as a single accept/reject bit performs identically to one reading randomly permuted verdicts (p = 0.57), so the harness discards the signal it already has. In contrast, zero-cost signals read off the loop's own dynamics (whether the answer changed, whether it oscillates, what the critique said) significantly predict both the benefit and the damage of another iteration (AUC gains of +0.137 and +0.097 over confidence). Refinement value is also front-loaded: only the first revision has significantly positive net value, later rounds are indistinguishable from zero, and by the fourth, harmful flips match successful corrections. Finally, signals of fixability transfer across task domains while signals of harm risk do not, an asymmetry that survives recomputing all features under a shared canonicalizer and suggests the safety half of a stopping rule needs per-domain calibration. The practical implication for agentic harness design is to instrument the loop rather than interrogate the agent.