Agreement With Humans Is Not a Property of the Judge
Diya Narayanan
Abstract
LLM judges are validated by correlating their scores with human labels, and that correlation is reported as a property of the judge. SummEval releases two annotator panels, three experts and five crowdworkers, who rated the same 1,600 summaries, and its authors report near-zero correlation between the two. This paper asks what that divergence does to a judge. One judge's scores correlate 0.43 - 0.61 with the expert panel and only 0.03 - 0.07 with the crowd panel, on identical items. The crowd panel is not noise. Its ratings carry significant variance tied to the source article ($\eta^2 = 0.11$ - $0.16$) and none tied to the system that produced the summary ($\eta^2 = 0.006$ - $0.012$, at the permutation null), where the experts reach $0.18$ - $0.34$: these annotators measured a real property consistently, just not the contrast the benchmark ranks. The shortfall is not a deficiency in the judge either, since a three-expert panel predicts the crowd labels no better, and scaling the judge from 0.5B to 32B parameters moves expert agreement by $+0.42$ and crowd agreement by $+0.01$. Both panels' aggregates are reliable ($R_3 = 0.79$, $R_5 = 0.83$), so a correction for annotator noise declares both targets sound; reliability is structurally blind to this failure. Since judge scores increasingly gate model comparisons, an agreement figure reported without naming the panel it was computed against cannot carry the conclusion it is asked to support.
Chat is not available.
Successful Page Load