BAJU: Label-Free Bayesian Auditing of LLM Judges Through Controlled Perturbations
Abstract
LLM-as-a-Judge enables scalable evaluation, but judge verdicts can be influenced by instability and bias. Existing audits often rely on correctness labels or report point estimates without uncertainty-aware deployment decisions, making them costly on unlabelled data and unable to determine whether observed risk is acceptably small or sufficiently supported by evidence. We introduce BAJU, a label-free, uncertainty-aware framework for auditing LLM judges for instability and bias before deployment. BAJU applies controlled, content-preserving perturbations, models perturbation-associated verdict changes with a Beta--Bernoulli posterior, and converts posterior exceedance risk relative to a pre-specified tolerance into auditable Pass/Warn/Fail decisions. Experiments on CALM's benchmark and a human-validated EvalBiasBench extension show that high point-estimate robustness does not guarantee certifiable reliability and that label-free evidence can usefully prioritise labelled follow-up. Under the stricter 5\% risk tolerance, BAJU flagged all four model--bias combinations classified as BIASED by the labelled benchmark for further investigation, while also exposing saturated baseline failure as an important limitation. These findings position BAJU as an inexpensive first-stage screen for unresolved or high-risk cases.