Towards Trustworthy Compact Judges: What Human Labels Can and Cannot Fix in Compact AI Judges
Abstract
Compact language models are increasingly attractive as evaluators because they can make large-scale evaluation substantially cheaper to run, repeat, and audit. But reducing the size of an evaluator raises a practical question: if a smaller model is less reliable, how much of that loss can be recovered with a small amount of human supervision? Answering this requires distinguishing two sources of disagreement with humans. An evaluator may produce scores on the wrong scale even when it ranks cases correctly, or it may fail to distinguish cases that humans judge differently. The first problem can plausibly be repaired with calibration; the second may reflect a limit on the model's underlying capability. We therefore ask whether human anchors can recover agreement lost by compact evaluators, and whether this recovery depends on what information the evaluator itself can represent. We find human supervision can improve a compact evaluator, but a judge can disagree with humans simply because its scores are too high or too low, and a small set of human examples can correct this mismatch. In our experiments, calibrating 72 SLMs on on 60 human-labelled examples, we find calibration therefore recovers much of the agreement that is lost because of score scaling. It does not, however, change which cases the model ranks above or below one another, particular for very small SLMs (<1B parameters). We see a related issue when models are compressed for deployment: 4-bit AWQ lowers the balanced accuracy of a 7B legal evaluator by 0.026, even though its self-consistency signal remains statistically unchanged. We propose a simple principle for lightweight evaluation: human labels can repair what a compact judge already knows how to distinguish, but they cannot upgrade its underlying capability. Labels should therefore be used to calibrate the judge's scale; missing distinctions require a more capable evaluator; and compression, confidence gates, and escalation should be validated on the target task rather than assumed to be safe. Compactness is valuable not because small judges are automatically trustworthy, but because they are cheap enough to calibrate, stress-test, rerun, and audit.