Active Preference Learning under Deployment Shift and Ties
Abstract
Preference labels are expensive in LLM post-training, motivating active methods that select the most informative comparisons. Many such methods assume a single, well-specified binary preference model. In practice, feedback sources, such as task domains or rater groups, may disagree, their proportions may differ between collection and deployment, and raters may report ties. When one reward model cannot represent all source preferences, changing where labels come from can change the learned reward model, even with unlimited data. We use a local bias–variance analysis to separate this persistent effect from estimation error. We derive the variance-minimizing source allocation among designs that use exact importance weights to preserve the deployment objective, and extend the approach to choose sources and response pairs jointly. For tied feedback, we characterize a transition in the Rao–Kupper model with a fixed tie parameter: when ties are sufficiently likely, comparisons with equal predicted rewards no longer maximize information about the reward difference. Estimating the tie parameter from the same comparisons can reduce the information available about the reward margin. Across three datasets with source shift, optimized importance weighting improves accuracy over weighting based on observed source frequencies, while log-loss results depend on calibration and the choice of regularization. Matching source counts to deployment proportions as closely as possible, without importance weighting, remains competitive. On datasets with ties, selecting comparisons near the predicted information peaks does not reliably outperform uniform sampling. Together, these findings show that active preference collection should account for both where feedback comes from and how preferences are expressed.