Sparse Repair in Reasoning Traces: A Structural View of Test-Time Inference
Zhuorui Zhang ⋅ Jianqi Yan ⋅ Shanshan Feng ⋅ Fan LI
Abstract
Inference-time reasoning improves when models spend more computation through repeated sampling, search, or verification. Existing scaling strategies usually allocate this computation across prompts, candidate answers, or whole trajectories. We study a different structure: how the repair value of additional inference is distributed within a single generated reasoning trace. Across mathematical, scientific, coding, and knowledge-intensive reasoning tasks, recoverable utility is sharply concentrated: in oracle replay analysis over coarse trace windows, the top-three windows capture 83--90\% of replay-recoverable utility. This concentration is not fully explained by position, length, or local uncertainty alone. Building on this finding, we introduce \textbf{LEVER}, a sparse local-intervention method for test-time reasoning that separates counterfactual discovery from deployable compute allocation. It uses local replay to estimate window-level recoverable utility, then distills this offline signal into a lightweight selector. At test time, the selector uses only online-safe trace features to decide which traces to activate and which windows receive additional rollouts. At a average 1.31$\times$ realized completion-token cost, LEVER improves average performance by about 4.2 points over base decoding. Across three base models and four task families, LEVER obtains the best result on 11 of 12 model--task pairs at the primary operating points, while using less compute than full-trajectory sampling baselines. These results suggest that effective test-time reasoning depends not only on how much computation is spent, but on whether it reaches the local commitments where repair value is highest.
Chat is not available.
Successful Page Load