AUDIT: Validating LLM Judges Through Controlled Error Injection
Md Mahmudur Rahman ⋅ Harsh Peswani ⋅ Elaine Wong ⋅ Dhruv Garg
Abstract
Validating LLM-as-judge evaluators by their agreement with human auditors is standard practice, yet no prior work reports how large the gap between agreement and actual error detection is in a deployed system. On a production agentic pipeline evaluated by a frontier LLM judge (46,342 human--machine comparisons, 35 rubric dimensions), we show agreement and recall are anti-correlated ($r = -0.672$), with median Cohen's $\kappa = 0.029$: the dimensions with the highest agreement are the worst at catching errors, so agreement, while useful for drift monitoring, cannot serve as a validation gate. The standard alternative is error injection: corrupt a known-correct output and measure whether the judge catches it. Prior meta-evaluation benchmarks assume injections are faithful; we measure this directly. We define infidelity $\delta$, the rate at which an injected item lacks the intended error or damages an unrelated dimension, and find $\delta = 0.807$ unfiltered, meaning four in five injections produce unusable test items. The boundary is structural rather than implementation-specific: $\delta = 0.019$ for factual-correctness dimensions versus $1.000$ for relevancy and completeness (0 of 204 attempts faithful). Correctness errors live in spans and can be corrupted in place; relevance and completeness are relations between query and response that cannot be broken by editing one side alone. For relation-valued dimensions, we recover auditing capability by changing what the response is measured against rather than what it contains: swapping the query while holding the response byte-identical raises the judge's flagging rate from 0.171 to 0.656 on 1{,}048 responses (label error 0.024), where injection produced zero usable items. Tool outputs carry nearly all remaining signal ($0.739 \to 0.196$ without them, $p < 10^{-5}$); without them the judge becomes permissive, not noisy, so the failure is invisible to monitoring that tracks only false-alarm rates. These findings establish three claims for practitioners: agreement is the wrong validation metric when errors are rare; injection without fidelity measurement is unreliable and inapplicable to relation-valued dimensions; and the most informative test requires no manufactured errors at all: if a change that preserves meaning still flips the verdict, the judge is unreliable by definition.
Chat is not available.
Successful Page Load