LLM Judges Punish Verbalized Evaluation Awareness
Abstract
Large language models increasingly recognize when they are being evaluated, and they sometimes say so in their reasoning. Verbalization of evaluation awareness is important because it allows us to monitor evaluation awareness and how it affects model behavior. In this work, we analyze the risk that using an LLM judge in training could suppress verbalization of evaluation awareness without eliminating it. Using a model organism, we show that LLM judges systematically disprefer rollouts with verbalized evaluation awareness across four judge models and four model character documents, including documents that never mention evaluation awareness. Further, we find that training using these judge preferences suppresses verbalization of evaluation awareness: verbalized awareness dropped from 44% to 9% after training. Finally, we show that even after evaluation awareness verbalization is almost completely suppressed, some evaluation-aware behaviors remain. More broadly, our findings show a mechanism through which training using LLM judges can shape model behavior via implicit judge biases.