When Does Long Thinking Help? A Discovery- Execution Framework for Mathematical Reasoning
Adib Hasan ⋅ Lay Jain ⋅ Thanic N Samin
Abstract
Prior work on test-time scaling gives mixed prescriptions: some results favor deeper reasoning, while others favor broader search. We study these behaviors by decomposing mathematical problem solving into strategy discovery and conditional execution, whose convolution determines end-to-end success under the framework's assumptions. Short unaided attempts and oracle-strategy interventions provide separate measurements of these components. Across four models on 35 new olympiad problems and 22 problems from IMO-ProofBench, the resulting model predicts performance under different depth--breadth allocations without observing their outcomes. The framework gives conditions favoring breadth or depth, with allocation equality under geometric discovery and execution saturated at $\varepsilon(1)=1$. Within Opus, accounting for execution improves predictions on problems where execution has not saturated in the first block. We also release the 35-problem benchmark with reference solutions and verified strategy sketches.
Chat is not available.
Successful Page Load