PriFT: Prior-Support Guided Token Reweighting for Supervised Fine-Tuning
Abstract
Supervised fine-tuning (SFT) is computationally efficient and broadly applicable, but often shows weaker generalization than reinforcement learning (RL). A key limitation is its off-policy objective: SFT fits fixed demonstrations token by token, including targets that may be poorly aligned with the model's pretrained distribution. Recent token-reweighted SFT methods address this issue by assigning larger training weights to tokens that better align with the model's predictive distribution, using statistics such as target-token probability or entropy. However, computing these statistics from the online model being fine-tuned makes token weights trajectory-dependent, as the model's distribution rapidly departs from the pretrained model and induces self-reinforcing reweighting dynamics. We propose PriFT, Prior-support guided Fine-Tuning, which derives token weights from a frozen pretrained reference to obtain a stable reweighting signal unaffected by online fine-tuning dynamics. This signal estimates prior support: the extent to which each target token is supported by the pretrained model before task-specific adaptation. Across multiple existing weighting and selection rules, replacing online statistics with pretrained statistics consistently improves performance. We introduce two instantiations: PriFT-prob, which uses pretrained target-token probability, and PriFT-mass, which selects tokens by relative support under the pretrained distribution. Extensive experiments on mathematical reasoning, code generation, and medical question answering show that PriFT achieves state-of-the-art SFT results and provides a better initialization for subsequent RL training.