Label-free detection and mitigation of judge bias
Abstract
Training language models against an LLM judge can fail invisibly: rather than improving response quality, models often learn to exploit superficial judge preferences that reward answers with bullet points, hedged phrasing, or excessive length. Because current bias detection methods require curated datasets, human labels, or extra inference costs, judges frequently remain unverified in standard pipelines. To address this, we propose a label-free method to detect and mitigate judge bias directly within the training pipeline. During preference tuning, models generate multiple responses per prompt. Because these paired responses typically share factual substance but differ in style, a robust judge should score them equally. Consequently, any recurring score advantage for a specific formatting feature provides a clear fingerprint of bias. We leverage this insight in two ways: as a post-hoc diagnostic to flag biased judges, and as an online training filter that intercepts and corrects unfairly high scores, preventing the model from learning cheap formatting tricks. We validate our approach by adding controlled bonuses to the scores from three open-weights judges. Our filtering mechanism successfully prevents models from adopting favored formats, and blind evaluations confirm these models perform just as well as those trained on unbiased judges. However, the method has limitations: it struggles to mitigate pure length bias without degrading quality, and the detection mechanism is sensitive to minor changes in phrasing.