Projected Distillation: On Learning to Solve Hard Problems from Close, Correct Targets
Abstract
On-policy learning is a cornerstone of post-training: by training a model on its own generations, it limits distribution shift and helps preserve existing capabilities. On hard problems, however, this signal thins out. On-policy RL provides no learning signal when every attempt fails and rewards are uniformly zero. On-policy distillation, which supervises the student's erroneous rollout token by token towards a teacher, is inefficient because feedback on later tokens remain conditioned on a mistaken prefix. Correct behavior can instead be imported off-policy from a gold or teacher solution, but such targets sit far from the model's generations and erode existing capabilities. We therefore view on-policyness as a continuum and ask where along it supervision should lie. We propose projected distillation, which trains on the correct solution closest to the model's failed attempt — a minimal correction refreshed online as the model evolves. Across two students (Qwen3-8B, OLMo-2-7B) and five tasks spanning math, code, science, and tool use, projected distillation achieves the best specialization–retention frontier outperforming the top baselines. Our ablations identify the key ingredients for near-on-policy external supervision: correction size, target-solution selection, correction refresh schedule, training objective and correction source. Theoretically, we formalize the continuum: our bound interpolates between DAgger-like linear-in-horizon regret for near-on-policy corrections and behavior-cloning-like quadratic regret for fully off-policy targets.