ExpRFT: Exponential Reward-Weighted Fine-Tuning for Offline RL in Multi-Turn Dialogue
Abstract
Offline reinforcement learning is a standard paradigm for fine-tuning Large Language Models (LLMs) on multi-turn dialogue, where on-policy rollouts and reward annotations are costly. Within this paradigm, reward-weighted supervised fine-tuning has emerged as an efficient approach, weighting each trajectory's log-likelihood by a reward-dependent coefficient. We show that existing reward-weighted SFT methods are specific members of a broader affine reward-weighted SFT class, and identify two structural limitations that hold for every member of this class: every centered member yields an unbounded-below loss, and no member admits a hyperparameter for attenuating reward noise. We introduce ExpRFT as a principled escape from this class — an exponential reweighting of the standardized advantage at a tunable temperature that restores a well-posed objective and provides a noise-control dial. ExpRFT parallels Advantage-Weighted Regression (AWR) in spirit, yet requires no per-action learned critic and adds no infrastructure cost beyond standard SFT. On six multi-turn dialogue benchmarks under reasoning and clarifying-question protocols, ExpRFT outperforms rejection-sampling, preference-based, value-based, and reward-weighted baselines. In the near-mean trajectory regime, where weight noise dominates the advantage signal, ExpRFT remains stable under injected reward noise while baselines significantly degrade.