Can We Trust Small Language Models as Judges?
Mingchuan Zou ⋅ Suchir Salhan ⋅ Paula Buttery
Abstract
Small language models (SLMs) make attractive judges: they are cheap enough to score every output, route uncertain cases, and calibrate to a domain with a handful of human labels. But practitioners often have to make deployment decisions from reliability numbers whose interpretation is unclear. We show that three common conclusions can be wrong: a judge can have higher agreement without better discrimination, improve under deferral without having useful uncertainty, and improve after calibration without benefiting reliably from the human labels. Across eight open-weight judges on essay and legal evaluation, we turn these ambiguities into practical audit tests. First, separate ranking from scale. Aligning judge scores to human mean and variance raises mean QWK from 0.246 to 0.433, but ordering remains capped at $\rho=0.444$; much of the apparent improvement is therefore score correction, not better judging. Second, benchmark deferral against the same review budget. Uncertainty-based escalation survives 500 matched-budget random policies in 37/40 legal conditions but 0/28 essay conditions, showing that a deferral gain does not by itself establish that the judge knows what it does not know. Third, treat human anchors as a source of uncertainty. We further show that high consistency can coexist with useless confidence and that near-constant outputs can appear reliable while failing to discriminate. We propose practioner recommendations: report ordering separately from calibration, compare escalation with same-budget random routing, report calibration performance across resampled anchor sets, and test whether confidence and consistency actually predict correctness. The practical lesson is simple: do not deploy an SLM judge because its reliability score looks good; deploy it when the specific capability that matters for the deployment has survived the corresponding counterfactual audit.
Chat is not available.
Successful Page Load