Who Should We Listen to More? Welfare-Aware Preference Acquisition for Pluralistic Alignment
Abstract
Pluralistic alignment requires learning how different people prefer models to respond. Model-generated preference labels offer scalable supervision, but their predictive quality varies across people, while acquiring observed feedback is costly. Under a shared, limited feedback budget, whose preferences should receive more learning effort? We develop Welfare-Weighted CABLE (W-CABLE), combining Cross-Model Acquisition with Brier Loss Estimation with designated target subgroup priority and sensitivity to current predictive disadvantage. Its exact marginal allocator maximizes a welfare objective over estimated predictive utilities. Across six Roleplay collections, we compare acquisition methods under equal-weight mean welfare and target-priority, inequality-sensitive welfare. At a 20\% query budget, W-CABLE improves Mean utility, Tail utility, and utility for a fixed, randomly selected 10\% target subgroup over matched baselines; across four main collections and both welfare specifications, Mean gains over the strongest baseline are 2.86--3.24\%. A parameter study characterizes tradeoffs among these outcomes. With training settings held fixed across acquisition methods, these gains translate into improved downstream alignment: under target-priority, inequality-sensitive welfare, W-CABLE with 20\% queried labels recovers 68.11--81.85\% of the Mean utility gap and 65.01--94.93\% of the Tail utility gap between training on all pseudo-labels and all ground-truth labels across Seen and Held-out personas.