Ultra-DPO: Reimagining LLM Alignment as a Semi-Supervised Task
Abstract
Direct Preference Optimization (DPO) offers a streamlined one-stage alternative to Reinforcement Learning from Human Feedback (RLHF). However, our theoretical analysis reveals a structural distinction between the two: DPO and RLHF optimize alignment on different data distributions with different supervision sources. Specifically, DPO learns from ground-truth human preferences but within the off-policy support of an offline dataset, while RLHF optimizes on the on-policy rollouts but learns from pseudo-labels produced by a reward proxy. This distinction explains how each method may fail: DPO alignment collapses when the policy drifts away from the offline dataset support, and RLHF alignment is biased due to the extrapolation error of the reward proxy on out-of-distribution rollouts. Effective alignment requires avoiding both failure modes, which neither method can do alone. We achieve this via a semi-supervised alignment objective: ground-truth preferences anchor alignment within the dataset support, and pseudo-labels on online rollouts extend it beyond. Applying DPO's reward reparameterization to this objective yields Ultra-DPO, a one-stage algorithm that augments direct human alignment with an online self-alignment mechanism. Across multiple benchmarks, Ultra-DPO consistently outperforms DPO, RLHF, and online preference learning methods.