Chain length is intrinsic time: process rewards and test-time compute as Bayesian information accounting
Akshay Balsubramani
Abstract
The length of a reasoning chain is the intrinsic time of a Bayesian sequential test. A chain monitored by a process-reward model stops when its posterior over correctness crosses a confidence boundary: Wald's sequential probability ratio test. Reaching confidence $1-\delta$ takes $\log(\delta^{-1})/K$ steps, where $K$ is the scorer's per-step relative-entropy gap. The empirical test-time-compute scaling laws follow from Pinsker's inequality applied to this identity, and the identity extends unchanged to multi-candidate scoring, decaying signals, adversarial rewards, calibration uncertainty and allocating one budget across many problems. The identity holds exactly on real reasoning chains from open models scored by public reward models. Read as an instrument, it prices a problem before solving it: a few-step probe reads off the compute cost about as well as measuring the rate exactly. An over-confident scorer stops uselessly early, and an affine recalibration of its log-odds restores the stop. Training the generator on the identity's dense reward raises the scorer's evidence rate and confidence at matched accuracy and more tokens, because the per-step reward telescopes to a terminal-confidence bonus. On this workload the probe is not worth its cost: an unprobed random order at equal total spend wins wherever the budget binds. Forcing an answer at the certified length loses it. And a calibration map carried from one dataset to another costs the stop up to half its verdict accuracy.
Chat is not available.
Successful Page Load