Estimating LLM Judge Error Rates and Outcome Prevalence with Limited or No Gold Labels
Abstract
LLM judges are widely used to label data at scale, and it is natural to query a judge several times and aggregate the results to address stochasticity. We investigate what those repeated queries are worth. Collecting 52,500 judge decisions from seven open-weight judges spanning 1B to 32B parameters, under three prompt conditions with ten independent samples per item, on a public benchmark carrying human labels, we find that ten repeated ratings carry the information of between 1.06 and 1.81 independent ratings. Judges repeat themselves rather than sampling afresh, so a practitioner who treats ten queries as ten observations will report an interval roughly 2.7 times too narrow. The consequences reach beyond aggregation. A latent class model built on a binomial likelihood, which encodes the same independence assumption, is badly miscalibrated. Across 42 cells its stated 95% interval for the judge's false acceptance rate contained the value computed from human labels zero times, and it underestimated total error in all 42. Replacing the binomial with a beta-binomial restores coverage to 0.90. We then show that recovery is governed by the gap between a judge's acceptance rate on genuinely correct and genuinely incorrect completions, a quantity estimable from a small labelled sample that predicts recovery almost monotonically. When the gap is small, a handful of gold labels restores accurate population estimates but not per-item classification. We release all ratings and code.