An LLM Judge for Subjective Behavior: Calibrate, Validate, Freeze, Then Optimize
Jules Roussel
Abstract
Subjective LLM behaviors are increasingly scored by LLM judges whose outputs drive optimization, yet whether a judge measures each criterion as human annotators meant it is rarely checked before its scores become a reward. We present a reproducible case study on MRBench tutoring dialogues (eight human-labeled criteria). A rubric judge (Claude Haiku 4.5) is developed on a dev split through human-reviewed rubric clarifications derived from judge--human disagreements, validated per criterion under a pre-registered rule, frozen with a hash-based drift guard, evaluated once on a held-out split, and re-scored three times for self-consistency. Two accepted revisions raised mean $\kappa$ from 0.47 to 0.57 (0.58 held-out); seven of eight criteria met the rule, and tutor tone was withheld from the objective. Optimizing a tutor prompt against the validated criteria raised the judge-measured macro pass rate on 60 held-out conversations from 59\% to 93\% (answer revealing 1\%$\rightarrow$89\%), stable across three runs. The pipeline turns a subjective behavior into criteria a judge is shown to measure, then improves the behavior on those criteria.
Chat is not available.
Successful Page Load