When Large Language Model Judges Meet Watermarks: Provenance Signals Distort Automated Evaluation
Osama Zafar ⋅ Alexander Nemecek ⋅ Erman Ayday
Abstract
Large language model (LLM) outputs are increasingly watermarked for provenance and increasingly evaluated by LLM judges, the same judges that rank leaderboards and generate preference labels for training. Watermarking rests on a neutrality claim: the embedded signal is statistically detectable but does not change how the text reads. That claim has been validated on human readers, never on the automated judges that now do much of the reading. Using 1,000 pre-LLM-era abstracts, six open-weight models, five watermark methods, and pairwise and n-way judging, we find that judges prefer controlled AI restyles over human originals and that watermarking lowers selection in all 20 cross-model target-judge pairs, by 11.1 percentage points on average. Self-judges largely stop favoring their own output, picking clean text 23.3\% of the time and watermarked text 8.7\%, and removing the matched clean copy narrows this gap but never closes it. Watermarking raises negative log-likelihood in all 125 judge-target-method cells, yet this distortion does not predict the cross-judge penalty ($r=-0.066$), ruling out a simple statistical-familiarity account. Watermark status is therefore a live confound in automated evaluation: a model that watermarks for compliance can lose comparisons for reasons unrelated to quality.
Chat is not available.
Successful Page Load