Audit the Rule and the Instrument: Two Checks Every Rubric-Evaluation Claim Should Report
Tianlang Chen
Abstract
A rubric-based evaluation report carries a point estimate and a nominal threshold. It rarely carries the two properties that decide whether its verdict means anything: whether the decision rule can resolve an effect at the scale it names, and whether the measurement pipeline returns the same answer twice. We propose reporting both, and demonstrate each on a separate study. In a preregistered gate asking whether a frozen, anchor-free pre-deployment rubric audit forecasts optimizer-induced proxy-anchor divergence, the audit does forecast out of cluster ($\rho = 0.39$) but adds nothing over a one-number summary of the same score matrix it already reads ($\rho = 0.45$; incremental $-0.0721$), failing its comparative condition by $+0.0045$ against a required $0.05$. By exact enumeration rather than simulation we then show that the rule producing that verdict could not reliably resolve its own $0.05$ threshold: at $K = 3$ held-out clusters the paired cluster bootstrap has 10 distinct multisets whose smallest atom exceeds the quantiles the rule uses, so every reported lower endpoint is the worst single cluster, and the rule's 80% detection crossing sits at a true margin of $0.20$-$0.32$. Separately, in a concurrent submission of ours, byte-identical reruns move 0.03%-2.30% of criterion verdicts across 12 judge-isolated pairs, the observed maximum movement those pairs induce in the downstream selection statistic is $0.6836$ pp, and candidate generation stopped reproducing in 3 of 18 generating runs with no mechanism identified. That study's own integrity audit returns a failing verdict, and Section 3 states the full scope of what we import from it. Neither defect is visible in any point estimate, and each is cheap to measure.
Chat is not available.
Successful Page Load