BiasTrojan: LLM Judgers Are Easily Distorted by Few Hundreds of Contrastive Biased Training Data
Abstract
Large Language Models (LLMs) are increasingly deployed as automated judges to scale supervision for data curation, reinforcement learning (RL), and agentic systems. While existing works have extensively explored bias in pretrained LLMs, the origins of such inherent biases remain largely untraced. We trace these biased tendencies to cognitively biased patterns (e.g., Authority, Bandwagon) latent in training corpora, and expose that such patterns are naturally prevalent in large-scale pretraining corpora yet remain entirely undetectable by existing data cleaning pipelines. To demonstrate that this underexplored threat can be deliberately exploited, we introduce BiasTrojan, a framework that concentrates and injects these naturally occurring bias patterns into training samples via context-aware bias cues and contrastive preference pairs, augmented with counterfeited reasoning chains for efficient injection. Experiments across several LLMs from 7B to 70B on human-preference and fact-related datasets show that mere hundreds of deliberately biased samples suffice to compromise LLMs into biased evaluators, overriding their factual knowledge. The injected biases generalize robustly out-of-domain and persist despite massive continual post-training. These findings reveal that this underexplored latent threat poses far greater risks than commonly recognized: biased LLM judges whose evaluations propagate irreversibly downstream, underscoring the critical need for bias-aware auditing and strict scrutiny of LLM training data.