CoTrek: Toward Scalable On-Policy Distillation for Long Chain-of-Thought Reasoning
Heng Zhang ⋅ Chengyu Zhou ⋅ Jiajun Wu ⋅ Estella Liu ⋅ Liheng Zhang ⋅ Yueqi Guo ⋅ Rui Liu ⋅ JiaHao Hong ⋅ Xuanxun Lian ⋅ Jinpeng Lu ⋅ Jin Huang
Abstract
On-policy distillation has emerged as an efficient paradigm for improving long chain-of-thought reasoning in large language models, where the student learns from its own rollouts while receiving dense teacher feedback on the states it actually visits. Prior work in this paradigm has focused on advancing \textit{continuation learning} through better teacher scoring, more stable optimization, or more learnable traces, yet largely overlooks a critical failure mode in long-horizon reasoning: once the student drifts at a deep prefix, the bottleneck is no longer how to \textit{continue} from the current path but how to \textit{recover} to an effective one. To address this gap, we propose \ourmethod, a recovery-centric framework for on-policy distillation in long chain-of-thought reasoning. \ourmethod operates on student-generated deep prefixes that have already drifted off track yet still admit successful teacher recovery. From each such prefix, it samples multiple teacher recoveries, extracts the short \textit{repair segment} they share, and trains the student to produce this repair before continuing the remaining reasoning on policy. This design shifts the distillation target from full-suffix imitation to focused path re-entry. To further support long-horizon learning, \ourmethod follows a depth curriculum that expands training depth according to the observed recovery gap, allowing the student to progressively extend its effective reasoning horizon. Extensive experiments on long chain-of-thought benchmarks demonstrate that \ourmethod delivers (I) \textbf{substantial performance gains}, improving strong on-policy distillation baselines by up to $6.84\%$ in final accuracy and $9.27\%$ in recovery success; and (II) \textbf{strong long-horizon robustness}, bringing $3.11\%$--$7.46\%$ gains on deeper prefix regimes and consistent pass@$k$ improvements under longer reasoning horizons.
Chat is not available.
Successful Page Load