Budget Boundary Effects in Test-Time Mathematical Reasoning
Abstract
A cumulative token cap can fall inside a mathematical derivation, forcing a test-time controller to choose between stopping at the cap (strict) and allowing the current attempt to finish (advisory). We measure this boundary choice with paired offline replays of 19,200 public traces: 120 AIME, BrUMO and HMMT problems and two archive configurations of one model. Candidate order and a 16-attempt cap are fixed, and answer selection is blind to reference answers and correctness labels. Three findings emerge. First, at the 4k cap, most advisory accuracy gains replace abstention with a correct answer; strict stopping pays for an unfinished prefix that the completed-only selector cannot use. Second, observed budget points reveal a different cost-accuracy picture: advisory 4k in the low configuration has higher accuracy than strict 8k at mean costs of 7,485 versus 7,940 tokens, while in high it approaches strict 32k accuracy using 59% of its mean tokens. Third, increased candidate coverage does not guarantee higher answer accuracy: a log-probability selector loses accuracy while coverage rises. Same-cap majority-accuracy differences shrink below 1.3 percentage points at 32k. We report coverage, vote transitions, parsing checks and realized cost separately. Budget curves for mathematical solvers should jointly state the cap, realized cost, eligible candidates, stopping rule and selector information.