Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Yuchen Cai ⋅ Ding Cao ⋅ Liang Lin ⋅ Chunxi Luo ⋅ Xin Xu ⋅ Kai Yang ⋅ Weijie Liu ⋅ Saiyong Yang ⋅ Tianxiang Zhao ⋅ Guangzhong Sun ⋅ Guiquan Liu ⋅ Junfeng Fang
Abstract
On-policy distillation (OPD) has emerged as an efficient post-training paradigm for large language models. However, existing studies largely attribute this advantage to denser and more stable supervision, while the parameter-level mechanisms underlying OPD's efficiency remain insufficiently understood. In this work, we argue that OPD's efficiency stems from a form of ``foresight'': it establishes a direct and stable path from the initial model to the final model early in training. This foresight manifests in two aspects. First, at the \textbf{Module-Allocation Level}, OPD identifies regions with low marginal utility and concentrates updates on modules that are more critical to reasoning. Second, at the \textbf{Update-Direction Level}, OPD exhibits stronger low-rank concentration, with its dominant subspace aligning closely with the final update subspace early in training. Motivated by this theory, we propose \textbf{EffOPD}, a plug-and-play acceleration method that speeds up OPD by searching for an effective step size and extrapolating along the current update direction. EffOPD requires no additional trainable modules or complex hyperparameter tuning, and achieves up to $3.5\times$ training acceleration while maintaining comparable final performance. Overall, our findings provide a parameter-dynamics perspective for understanding the efficiency of OPD and offer practical insights for designing more efficient post-training methods for large language models. Our code is available at: https://anonymous.4open.science/r/EffOPD-7C58.
Chat is not available.
Successful Page Load