Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Abstract
LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference (NLI) datasets with 100 human annotations per item, we find that the 9 judges effectively provide only about 2 independent votes’ worth of information (Kish neff = 2.18 on MNLI). The panel’s accuracy falls 7–22 percentage points short of the agreement-matched Condorcet prediction (7–16 points under a plain independent-voting simulation), and the panel never meaningfully outperforms the best single judge (maximum lift +0.4pp, within noise). Established stable aggregation methods close at most 11% of this gap even with access to the correct answers. The deficit is stable (neff ≈1.9–2.5) across prompt variants, temperatures, chain-of-thought, and a pairwise preference task (RewardBench). We distill the findings into five concrete diagnostics for evaluation-panel builders.