Auditing a Conflicted Judge: Rater-Free Verification of an LLM-Scored Forensic Reliability Benchmark
Abstract
Benchmarks for high-stakes domains increasingly use an LLM as the scorer. When the scorer is drawn from the same family as the systems it scores, the resulting leaderboard is confounded in a way no amount of careful prompting resolves. We report a benchmark that fell into exactly this trap and the audit we used to determine what, if anything, survived it. We construct a 30-case synthetic Android forensics benchmark scoring the evidentiary reliability of model-generated findings (evidence support, provenance, observation/inference separation, attribution justification, and reproducibility) rather than their textual similarity to a reference answer. Four models (three Claude models and the open-weight qwen2.5:7b) produce 120 responses containing 830 scored claims. The primary annotator is Claude Sonnet 5: the same model as one system under test and the same family as two more. Rather than disclose this and proceed, we audit the annotator with two independent instruments. A cross-family LLM rater agrees poorly in absolute terms (Krippendorff's alpha = 0.052) and is significantly stricter (p = 0.0004). We then build a rater-free deterministic verifier: four checks computed from case ground truth by string and regex operations with no model in the loop, hand-validated to 100% precision over every flag it raises. The verifier recovers the annotator's top tier exactly, with both frontier Claude models committing zero violations against 9 for the two weaker models (Fisher p = 0.0046). It independently confirms the annotator's most suspicious result, a perfect attribution-dimension ceiling, by finding zero person-attribution violations across all 120 responses. It also reverses the annotator on the lower-tier ordering, ranking Haiku above qwen2.5:7b where the annotator ranked it below. That disagreement runs opposite to self-preference: the conflicted annotator was harsher, not kinder, to the same-family model. We argue that partial rater-free verification should be a routine component of LLM-judged benchmarks, and that its main value is not validating the judge but marking precisely which conclusions do not need it.