Behind Aggregate Scores: What Systematically Varying Arena-Hard's LLM Judge Reveals
Abstract
Different LLM-judge setups can yield nearly identical aggregate scores while changing many prompt-level outcomes. The gap matters for model rankings, preference labels and safety review. We test this on 500 Arena-Hard v2.0 prompts, holding responses fixed and varying only the judge. GPT-4.1 judges the fixed answers under all eight combinations of three instruction choices: Arena-Hard's original criteria or a rubric prioritising factual and technical accuracy; reasoning before the verdict or the verdict label alone; no task description or a coding-assistance description. We substitute Claude Sonnet 4.5 or Gemini 2.5 Flash under the original instructions, and we also repeat the baseline specification unchanged. The repeat is the reference comparison. It moves the aggregate score by 0.10 points while changing the winner-or-tie outcome for 20.0% of prompts, so the two levels come apart before any specification changes. Judge changes add to that floor rather than creating it. Label-only output changes 22.0 percentage points more prompts than the repeat and lowers the score by 3.20 points; Gemini changes 19.0 points more while moving the score 0.18 points; a coding-assistance description is indistinguishable from the repeat. Consequential evaluations should test a small, prespecified set of alternatives plus one unchanged repeat, and report changes at both levels against that floor.