On the Sparsity of Direct Preference Optimization: Weight Disentanglement in the NTK Regime
Kensuke Sasaki ⋅ Issei Sato
Abstract
Direct Preference Optimization (DPO) has been widely adopted for aligning language models with human preferences. Recent empirical work has revealed that DPO induces remarkably sparse parameter updates, yet no theoretical explanation for this phenomenon has been offered. We analyze DPO's optimization dynamics under a neural tangent kernel linearization and show that the logistic structure of the DPO loss gives rise to an implicit maximum-margin bias. Building on this observation, we introduce the concept of dual sparsity: the DPO update direction is a sparse linear combination (data-level sparsity, driven by support vectors) of sparse gradient feature vectors (parameter-level sparsity, arising from gradient cancellation between preferred and dispreferred responses). We formalize this structure in our main theorem and present a framework relating it to weight disentanglement---a condition favorable for model merging. Experiments on GPT-2 Small and Medium with the Anthropic HH-RLHF and UltraFeedback datasets confirm both sources of sparsity and show that DPO reduces cross-task interference by up to $35 \times$ compared to supervised fine-tuning.
Chat is not available.
Successful Page Load