Escaping the Token-by-Token Learning Trap in On-Policy Distillation for Mathematical Reasoning
Abstract
To improve LLMs' mathematical reasoning ability, On-Policy Distillation (OPD) trains a student on its own solution trajectories under dense teacher supervision. However, standard OPD applies this supervision token by token. It may correct the student's immediate prediction without redirecting the mathematical reasoning that follows—a failure we call the token-by-token learning trap. Token loss also provides a noisy starting point: about 30% of high-loss positions induce little near-future trajectory divergence and often reflect surface-form differences rather than reasoning errors. We propose Trajectory-Aware OPD (TOPD), which compares short teacher and student continuations from the same prefix, uses optimal transport to identify consequential divergence, and distributes teacher guidance across the future window. Distilling Qwen3-30B-A3B into Qwen3-4B, TOPD improves average accuracy over standard OPD from 47.8% to 52.2% on AIME24, AIME25, and HMMT25-Feb. These results show that mathematical reasoning distillation benefits from supervising how a solution trajectory develops, rather than only which token comes next.