Don't Discard Your Rollouts: Reusing Teacher RL Traces for Student Distillation
Abstract
RL post-training with methods like PPO and GRPO produce a large rollout archive, which each contain correctness-labeled completions, correct/incorrect completions for the same prompt, teacher log-probabilities, and group pass rates. Canonical distillation methods, however, ignore these traces and train based upon new data from the converged teacher. This discards the contrastive signal that helped the teacher and the metadata already produced during teacher RL, even though the teacher RL run typically generates many more tokens than the later distillation pass. However, reusing rollouts naively is non-trivial: the archive mixes early and late teacher policies, can be arbitrarily off-policy for the student, and contains many low-quality traces, and naive SFT and DPO on these rollouts underperforms the cannonical SFT baseline. In this paper, we show that rollout reuse becomes effective when it is treated as a data-selection problem. We use two sets of information from rollouts: per-prompt pass rate, which defines an easy-to-hard curriculum and separates all-correct prompts for SFT from mixed-success prompts for DPO; and per-rollout teacher likelihood, which acts as an in-distribution proxy for student-compatible traces. Across two GRPO teachers (4B and 8B) and students from multiple families, the curated rollout datasets DSFT^{allp,top1} and DDPO^{curr,top1} match or beat canonical distillation without an extra teacher generation pass. On MATH500, DDPO^{curr,top1} improves over Dsynth by up to 4.0 percentage points on strong students, while DSFT^{allp,top1} → DDPO^{curr,top1} matches weaker students where DPO is unstable.