When Checks and Verdicts Disagree: Auditing Aggregation-Contract Fidelity in a Prompt-Structured LLM Judge Family
Abstract
Analyses of LLM judges increasingly run on their structured outputs, reading per-criterion checks as a decomposition of the verdict. That reading assumes something rarely checked: that the stored verdict really is what the judge's own per-check fields imply, under the enumeration the prompt declares and the aggregation rule it induces. We make the assumption executable: attempt to recompute every verdict from the released per-check returns. We run this on 29,352 record–configuration cases — 1,223 distinct records clustered in 215 source papers, each judged under twenty-four configurations of one eleven-check prompt family (twelve backbones, two prompt variants). The roster retains no raw response text, so what is audited is the parsed-output judge pipeline — the parsed-and-stored field set, not the token stream that produced it. Observed non-derivability varies widely across the roster — median 3.15%, four configurations with no observed non-derivability — and a single violation rate hides both the size and the kind of the failure: on the highest-failure configuration 22.81% of stored verdicts are non-derivable from their own checks, while a mismatch-only tally over all 1,223 records reports 11.12%, because 143 outputs cannot be aggregated at all. A post-hoc, non-randomized same-window diagnostic illustrates that the observed rate is not determined by rubric content alone: requesting the checks before the verdict was associated with conditional inconsistency dropping from 10.48% to 0.08% (llama-3.3). The twenty-four configurations share one task, one schema and one corpus: this is a prompt-family case study, not evidence about judges in general.