Measuring the Measurer: Reliability of LLM Judge Scoring in Petri
Adeline Kassler ⋅ Axel Ahlqvist ⋅ Juan-Pablo Rivera ⋅ John Hughes
Abstract
Automated safety evaluations hinge on the reliability of LLM judges. Petri, an open-source adversarial evaluation framework used in official system cards, scores audit transcripts on 38 behavioral dimensions in a single judge call. Measuring this scoring across four corpora (including a synthetic calibration corpus with known planted behaviors), up to six judge models, and two scoring modes, we find that configuration choices usually treated as implementation details largely determine what the evaluation reports. Judge models disagree with each other more than with themselves (per-judge ratios $1.8$--$10.9\times$, median $2.8\times$) and the disagreement persists at temperature 0; the default single-pass mode compresses mid-range scores by up to 4.3 points and is sensitive even to the order in which dimensions appear in the prompt. Remedies fail in characteristic ways: per-dimension scoring detects more planted behaviors ($69.3\%$ vs.\ $62.3\%$) but at a higher false-positive rate ($28.9\%$ vs.\ $18.7\%$) and nearly double the inter-judge disagreement; detailed rubrics do not improve detection accuracy; ensemble stability is an arithmetic consequence of averaging. The patterns are not artifacts of one rubric or harness: between-judge disagreement exceeds within-judge retest noise under all seven dimension lists we test and on two external instruments scored with their own shipped prompts, and in a downstream auditing pipeline the choice of judge alone moves the published violation rate for a fixed target from $13\%$ to $32\%$. Judge configuration is part of the measurement instrument and should be designed, disclosed, and audited as such. We release our corpora, scoring data, and rubrics.
Chat is not available.
Successful Page Load