Do LLM Judges Measure What We Mean? Auditing Criterion Alignment and Metadata Dependence in RAG Evaluation
Konrad Wojtasik ⋅ Marcin Oleksy ⋅ Maciej Piasecki
Abstract
Whether an LLM judge can be trusted is a property of the entire evaluation protocol, the criteria it is given, the evidence it sees, the form it fills in, and how its scores are interpreted, and not of the model alone. We audit that protocol on 386 answers to 65 Polish public-sector RAG questions, scored by six LLM judges. Replacing a generic criterion with the human annotation policy raises the mean quadratic weighted kappa by $0.170$ for correctness and $0.176$ for conciseness, and every one of the six judges improves. But judges agreeing with each other is not the same as judges being correct. One prompt reaches ordinal Krippendorff's $\alpha=0.618$ across the panel and only $QWK=0.101$ against the experts; it awards the maximum score in 85\% of decisions. We performed ablation studies to determine which components affect proper alignment with human annotations by removing the reference gold answer, helper binary questions, and document relevance annotations. The core results indicate that clear human annotation guidelines are crucial before conducting LLM-as-a-judge evaluations.
Chat is not available.
Successful Page Load