Confidence-Calibrated Inference Expansion for Evaluator-Guided Test-Time Reasoning
Abstract
Evaluator-guided reasoning systems face a local control problem: given a current answer and evaluator evidence, should the system trust, fallback, or buy more inference? We formulate this as a compute-aware three-action controller over trust, fallback, and expand, where each action is scored by predicted eventual success minus explicit compute cost. The central quantity is action-specific confidence: the success probability induced by trusting now, falling back now, or expanding once and continuing. We show that local action selection is stable when confidence error is smaller than the compute-adjusted action gap, and that regret accumulates only along the realized expansion path. Guided by this analysis, we learn one confidence head per action, calibrate the heads on held-out data, and penalize locally uncertain actions. Across verifier-guided width expansion on MATH and AIME25 and judge-guided revision expansion on a deterministic-answer General Reasoning benchmark, calibrated control improves the utility--compute frontier relative to strong fixed and adaptive baselines. The same pattern persists under policy-induced expansion targets, same-feature budget routers, fallback removal, stronger learned controllers trained on the same features, and moderate evaluator mismatch after refitting and recalibration. When compute is free, high-budget policies can remain competitive; the gain here is compute-efficient allocation, not a uniformly stronger reasoner.