Commit Then Explore: Reasoning Models Diversify, Not Converge
Ben Jenkins ⋅ Mihaela Cardei
Abstract
Test-time compute scaling, where language models "think longer" via extended chain-of-thought (CoT), is the dominant paradigm for improving reasoning. Prevailing intuition holds that extended reasoning *converges* toward an answer; we show the opposite. Across three reasoning-RL families (DeepSeek-R1, Qwen-QwQ, Llama-3.1-Nemotron) and three benchmarks (GSM8K, MATH, GPQA), step-to-step SAE feature turnover and entropy *rise* monotonically across the chain (Cohen's $d = 0.61$–$0.68$, $p < 10^{-3}$). Yet the final answer is already linearly decodable from the residual stream within the first $\sim 20\%$ of the chain. We call this pattern **commit then explore**: the answer becomes stably decodable early, and the bulk of test-time compute is then spent exploring *around* an already-decodable answer. Targeted ablation shows post-commit features have a $2$–$6\times$ smaller teacher-forced log-probability effect (matched-mass control: $1.5$–$4.6\times$) and a $2$–$3\times$ smaller free-regeneration answer-flip rate than pre-commit features. A terminal-reward RLVR credit-assignment account predicts this asymmetry, and the dynamic is absent in base models: Qwen-2.5 and Llama-3.1 bases fail our detector at every layer, and a held-out quantitative prediction derived from R1 generalises to Qwen-QwQ at Pearson $r = 0.96$. Leveraging the structure, we build a Reasoning Monitor that saves $26$–$34\%$ of inference compute at $\geq 92\%$ accuracy retention, dominating four adaptive-compute baselines. Extended chain-of-thought, under this view, is not the computation of an answer; it is the exploration of its neighborhood.
Chat is not available.
Successful Page Load