When the Grader Is Fooled: Measuring How Often LLM Judges Accept Plausible-Wrong Answers Across Six Model Families
Abstract
Using a large language model to grade another model's output has become routine, and increasingly the grader's verdict is simply believed. We audit the grader instead. On a controlled pool of items with objective keys (arithmetic word problems, Python functions checked by unit tests, and factual questions) we ask twelve judges drawn from six vendor families (Anthropic, Amazon, Alibaba, Meta, OpenAI, Google) to grade candidate answers, half of which are wrong by construction. Every item is graded three times at temperature zero; in all we collected and stored 60,801 verdicts. The failures are not random. Judges almost never accept a blatantly wrong answer (at most 0.4% for all twelve judges on a domain-matched control), yet wrong answers that look right get through at rates that split the tested snapshots into two camps: 1.4–3.3% for the OpenAI (GPT-OSS), Anthropic (Claude), and Alibaba (Qwen) judges we tested, versus 19.7%, 25.2%, and 34.6% for the Amazon (Nova), Google (Gemma-3), and Meta (Llama) judges, and up to 62% on off-by-a-small-amount arithmetic. A between-family variance decomposition shows this is a leniency axis: after removing per-item structure, the between-family component (0.0062) is an order of magnitude smaller than the per-item orthogonal residual (0.068); whether the residual shared component reflects a common blind spot or item-level confounds is not identifiable in this design; what is identifiable is that the permissive families' errors correlate with one another. A strict "derive the answer yourself first" rubric narrows the gap but does not close it, and it is not free: for the permissive judges it buys fewer false accepts at the cost of up to 16 points of false rejection on correct answers. The harness, the raw verdict cache, and analysis code that regenerates every number offline are available at https://anonymous.4open.science/r/grader-fooled-2026; a public repository accompanies the camera-ready.