Can We Trust Adversarial Filtering? When the Judge Curates the Benchmark
Abstract
Adversarial filtering shapes a benchmark by discarding questions according to model performance, to reduce annotation artifacts, preserve headroom, and automate benchmark design. Filtering judges—the LLM judges whose labels drive this loop—are typically admitted by how well their labels agree with human labels on fixed responses. We ask whether judges that pass this check interchangeably go on to build the same benchmark. They do not: five judges within a 4.4-point band of macro-F1 against human experts produce five different benchmarks, the questions they disagree on are among the most discriminative in the pool—exactly the questions a benchmark can least afford to lose—and the same judge-dependence reappears on published labels from eleven other data collections, with human as well as LLM judges. Our controlled replay runs a complete adversarial filter over 1,136 questions written by human experts in seven African languages, holding questions, model responses, prompt, references, and thresholds fixed and changing only the judge. The five judges disagree on the membership of 157 questions (13.8%), and the resulting sets shift every evaluated model's score, including which model attains the highest point estimate. Those contested questions separate the eight evaluated models roughly three times more sharply than questions all five judges accept (43.2 versus 13.4 percentage points), a gap that persists after matching language and observed difficulty. On the external collections, including SummEval, WMT MQM, JuStRank, and WildBench, the strongest cases survive multiple-comparison correction. Because response-level agreement with human labels does not stabilize benchmark construction, builders should replay their filter across several judges and refer contested questions to human experts rather than discarding them automatically.