An Axiomatic Analysis of DPO and NLHF as Reference-Dependent Probabilistic Voting Rules
Abstract
Post-training methods for large language models (LLMs) implicitly aggregate diverse human feedback in ways that can be studied explicitly using social choice theory. In this paper, we adopt the perspective that for each post-training method, there is a corresponding reference-dependent probabilistic voting rule. Taking Direct Preference Optimization (DPO) and Nash Learning from Human Feedback (NLHF) as our case studies, we study axiomatic properties of the associated voting rules. First, we show that NLHF violates one of the central axioms of social choice, namely monotonicity, if and only if its KL regularization parameter is below a bound, whereas DPO satisfies monotonicity regardless of its regularization parameter. Second, we show that while DPO is often associated with Borda-style preference aggregation, when DPO is regarded as a probabilistic voting rule, it violates two variable population axioms that the standard probabilistic version of Borda satisfies. Third, we show that unlike DPO, NLHF violates a commutativity axiom from Bayesian epistemology, which has practical implications for staged post-training. Finally, we conduct an empirical study of the frequency and magnitude of axiom violations in the CoVal dataset using real base models.