Finetuning with Sampling: Make SFT Generalize, Not Forget
Abstract
Introducing new capabilities to frontier models has long been the goal for posttraining, which relies predominantly on supervised finetuning (SFT) and reinforcement learning (RL) to achieve this. Prevailing wisdom dictates that on-policy RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and forgetting. At the same time, SFT enables learning from inherently off-policy expert data, whereas RL must rely on a model's ability to generate positive signal on a new task. In our work, we seek to bridge the strength of on-policy learning with the information signal of off-policy expert traces. To this end, we introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that transforms off-policy traces into more on-policy ones given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and fact learning, we demonstrate that simple SFT on these boosted off-policy traces can generalize better and forget less than on-policy counterparts. More generally, our approach outlines a principled methodology for model-specific data refinement, suggesting broader utility as a plug-and-play component throughout the post-training pipeline.