Correct, Route, Calibrate: Efficient Preference Optimization from Noisy, Heterogeneous Human Feedback
Abstract
Human labels are important supervisory signals for preference-based reinforcement learning, but their collection is often constrained by practical considerations that standard training objectives tend to ignore. For instance, Direct Preference Optimization (DPO) usually treats observed preference labels as clean, even though real annotation pipelines rely on heterogeneous annotators whose reliability varies across individuals and with the difficulty of each comparison. This setting introduces three challenges: (1) Noisy labels introduce systematic bias into the preference optimization objective and need to be rigorously corrected. (2) The allocation of pairwise comparisons to different annotators is a design problem that directly affects the statistical efficiency of the training pipeline. (3) Non-random routing creates selection bias in annotator evaluation, complicating downstream decisions such as compensation. We address these challenges within a unified framework for DPO that combines annotator-aware posterior correction with information-guided annotator routing. Posterior correction leverages an instance-dependent noise model to infer latent ground-truth preferences from noisy labels, mitigating bias in the DPO objective. Annotator routing is formulated as an optimal experimental design problem that allocates comparisons to annotators according to their expected statistical information under a fixed annotation budget. To support reliable annotator evaluation under non-random routing, we introduce a doubly robust estimator that corrects for selection bias. Semi-synthetic experiments on the UltraFeedback dataset show that our approach improves ground-truth preference recovery and downstream AlpacaEval 2 performance over baselines, with consistent gains across multiple open-source LLM model architectures. Annotator reliability-estimation experiments further show that the doubly robust estimator substantially reduces mean squared error relative to other baselines, demonstrating its value for annotator evaluation.