When Scores Conflict with Preferences: Calibrated Drift Control for Heterogeneous DPO
Ping Liu ⋅ Yan Yan
Abstract
Direct Preference Optimization (DPO) is increasingly used in settings where pairwise preference signals coexist with pointwise criteria such as response length, answer correctness, or safety scores. When the two disagree, standard DPO does not explicitly distinguish pairs that improve the pointwise criterion from those that worsen it. We show that this mismatch induces a systematic failure mode, which we call **constraint drift**: updates on disagreement pairs tend to increase pointwise violation, while updates on aligned pairs tend to repair it, with the overall drift governed by their balance. Under a first-order approximation, this balance yields a closed-form pre-alignment diagnostic, $\rho^{*}_\mathrm{eq}$, computable from dataset statistics alone, that predicts whether the data is self-correcting or requires intervention before alignment begins. Building on this diagnostic, we propose **Conflict-aware DPO (CoDPO)**, a calibrated modification to the DPO loss with two complementary components: a coarse per-sample control for primary drift regulation and a fine adaptive margin tilting for residual correction during training. Across six domains, drift grows monotonically with conflict exposure, and $\rho^{*}_\mathrm{eq}$ tracks the empirical drift transition, including self-correcting settings where intervention is unnecessary. Across four base-model architectures and multiple preference losses, CoDPO reduces drift while preserving preference learning. Compared with hard filtering and recent constrained-alignment and multi-objective baselines, CoDPO achieves a more favorable trade-off between drift control and preference performance.
Chat is not available.
Successful Page Load