When the Judge Selects the Data: An Outcome Audit of LLM Answer-Preservation Screens
Abstract
Prompt-robustness benchmarks use large language models (LLMs) to remove perturbations that may change task meaning. Because these judges select the evaluation sample, their errors can alter benchmark composition. We audit two complementary settings. Screen~1 audits a published filter on typos and stopword deletion. Screen~2 broadens the evaluation to six benchmarks and three paraphrase types. For each screen, we ask whether answering models fail more often on rejected than accepted perturbations. Across 1,186 perturbation instances from six multiple-choice benchmarks, seven judges tested under subsets of eight configurations select sharply different rejection sets. The published Screen~1 pair shares only 10\% of rejections. A published rule removes nearly half of Screen~1, yet rejected and retained prompts fail at almost the same rate. Small, shared rejection sets are more informative: a three-judge consensus removes 5\% and identifies a substantially more failure-prone subset. On Screen~2, filtering removes up to 20\% of items but changes every robustness estimate by at most 2.7 percentage points. Overall, screen design strongly shapes the selected data, while aggressive filtering can discard examples without improving benchmark signal. We recommend reporting benchmarking results with and without filtering.