The Post-Training Dilemma: Why We Should Rethink the Sequential SFT-RL Paradigm
Xueyan Niu ⋅ Bo Bai ⋅ Wei Han ⋅ Weixi Zhang
Abstract
This position paper argues that the sequential supervised fine-tuning followed by reinforcement learning (SFT-then-RL) pipeline---the de facto paradigm for large language model (LLM) post-training---has a fundamental property that cannot be fixed by hyperparameter tuning: SFT and RL optimize over nearly orthogonal parameter subspaces, so RL inevitably degrades the capabilities acquired during the preceding SFT phase. Under the PL condition, we show that this degradation scales linearly with the number of model parameters. Although the KL penalty can constrain distributional shift, it cannot correct the underlying gradient misalignment between the two objectives. Experiments on Qwen3-0.6B across four tasks confirm the theoretical predictions: on MATH-500, Alpaca, HH-RLHF, and SST-2, the SFT loss increases after the RL phase. We identify the underlying mechanism as persistent near-orthogonality between the gradients of the two objectives ( $\cos(g_\mathrm{SFT}, g_\mathrm{RL}) \approx 0$), paired with opposing curvature dynamics: SFT flattens the loss landscape while RL re-sharpens it. Based on our analysis, we propose three promising research directions: (1) joint optimization of the two objectives with gradient surgery to resolve conflicting updates; (2) landscape-aware scheduling that terminates RL training precisely at the theoretically optimal point; and (3) characterizing the model scale threshold above which SFT becomes unnecessary.
Chat is not available.
Successful Page Load