Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level and Invisible in a Six-Model Leaderboard
Abstract
Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys, on two corpora that put predefined annotator populations on the same items. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, the normal and expert-designated pools assign a different majority label to 23.6% and name the opposite winner on 9.2%; on the 246 comparably unanimous MT-Bench cells, benchmark authors and recruited experts differ on 30.5% and reverse on 8.5%. Yet on both corpora the resulting six-model orderings are identical: Kendall tau = 1.00, zero models displaced, even though the win-rate vectors are not. That invariance is weaker evidence than it looks, and the weakness is measurable on this corpus rather than extrapolated. An item-level bootstrap of the real corpus displaces at least one model in 27.9% of resamples on the full corpus and 34.7% on the convention-free subset the headline uses, almost all of it at the one adjacent pair sitting 0.8pp apart, which the two golds order differently in 22.6% of resamples. These are resampling frequencies on this corpus, not a general probability of leaderboard change. The observed zero is the modal outcome rather than a reliable one, not a property of aggregation, and it is a statement about this leaderboard's spacing rather than about leaderboards. The per-item consumer we test, three local LLM judges, leans toward the normal pool by 1.8-3.5pp after reweighting to the corpus when scored against tie-broken pool majorities, with one of the three distinguishable from zero. That sign depends on the tie-break: one judge reverses under the leaderboard's own rule, and a per-annotator score that needs none leaves one judge with a lean distinguishable from zero. We separate the two regimes. Within MultiPref's expert-designated pool the two annotators give different coarse labels on 50.1% of comparisons, evidence against the zero intra-group variability its routing model assumes and distinct from the between-pool contrast its authors left for future work. Code, results and per-call judge outputs: https://github.com/anik-jha/whose-gold.