Margin Dynamics for Large Language Model Alignment
Xingzi Xu ⋅ Saygin Seyfioglu ⋅ Karim Bouyarmane
Abstract
Practitioners choose among LLM alignment methods largely by trial and error. We provide a theoretical foundation: each pairwise method defines a vector field on the probability simplex that governs how the model redistributes mass between preferred and dispreferred completions. Projecting onto a single preferred/dispreferred pair reduces this to a one-dimensional margin ODE with a closed-form solution. These solutions reveal that DPO margins grow as $O(\log t)$ while IPO margins converge exponentially to $1/(2\tau)$; that label noise induces a qualitative phase transition in DPO (unbounded $\to$ bounded) but affects IPO only quantitatively; and that gradient concentration under asymmetric weighting governs trainability at scale. We confirm all three predictions across five model families spanning four benchmarks. The closed-form solutions also enable principled method design: we prove that adding \emph{any} label smoothing $\varepsilon > 0$ to SimPO creates a finite margin attractor at $\Delta^* = \beta^{-1}\log((1{-}\varepsilon)/\varepsilon)$, while $\varepsilon{=}0$ (standard SimPO) has no attractor. We test this across three model scales (3.8B--9B) and confirm it: the theory-derived configuration (TPO-S, $\varepsilon{=}0.05$) achieves the highest mean IFEval score on all three models.
Chat is not available.
Successful Page Load