The Errors That Matter: Consequence-Aware Evaluation of LLM Judges
Vishal Srivastava
Abstract
LLM-as-a-judge pipelines increasingly determine which model outputs are accepted, rejected, or escalated, yet human verification capacity is limited. A natural strategy is to review judgments that appear most likely to be wrong. We ask a different question: when judge errors have heterogeneous consequences, does error-likelihood-based review prioritize the errors that matter most? We introduce Materiality-Bench, a frozen benchmark of 150 controlled synthetic policy-reasoning cases in 50 matched groups. Crucially, we separate the ex-ante materiality signal available to an allocator from the evaluation-only realized consequence of an observed wrong decision, with both label sets human-reviewed and frozen in sequence before final outcome analysis. For GPT-5 nano, 21 of 22 reference-label errors carry nonzero realized consequence, yet 74.1% of realized consequence mass lies in the lower half of cases ranked by cross-fitted estimated error probability ($D_{50}=0.741$, group-bootstrap 95% CI [0.548, 1.000]). At a 20% review budget, error-likelihood review leaves 92.6% of realized consequence unresolved, compared with 59.3% under consequence-aware allocation. Integrated over budgets from 0–50%, the advantage is $AUC_{\Delta}=0.171$ (95% CI [0.079, 0.253]). Gemini shows a smaller positive allocation effect but does not independently satisfy the frozen decoupling criterion; Qwen makes no reference-label errors and is therefore non-informative for consequence allocation. Materiality-only allocation is competitive with the full probability-times-materiality score, and a within-group permutation control does not establish that fine-grained case-specific materiality placement is necessary. The supported conclusion is deliberately narrow: when judge confidence provides little usable triage resolution, consequence information can improve how a limited human-review budget is allocated.
Chat is not available.
Successful Page Load