Codebook-Driven Correction of LLM Evaluators: A Counter-Narrative Case Study
Abstract
When an LLM judge disagrees with a strong human consensus, the disagreements can follow a pattern. In a case study on counter-narrative evaluation, we find that they cluster into recurring, interpretable defect types. The judge rewards responses that dodge the actual claim, forgives AI refusal boilerplate, and lets whataboutism pass unpenalised. We present codebook-driven evaluator correction, a task-agnostic framework that turns these patterns into a diagnostic and corrective tool. From free-form critiques of every dispreferred output in a benchmark of high-consensus pairwise comparisons, the framework induces a compact codebook of the defect types that drive human preference, projects any judge's disagreements onto it as an error profile, fine-tunes the judge on type-targeted contrastive pairs, and reads the update as a redistribution of error mass across the same types. On a 559-comparison benchmark, the induction yields an eleven-type codebook that independent annotators apply with substantial reliability (Fleiss' kappa = 0.755). Correcting a Qwen3.5-9B judge recovers 70% of its baseline errors while preserving 94.2% of its baseline-correct decisions. Read type by type, the nine-point gain is a real contraction of the judge's error set rather than a reshuffle, and it exposes one localised regression that the aggregate would have hidden.