Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate
Abstract
Text watermarking helps identify AI-generated content, but its effect on factual reliability remains underexplored. In this paper, we study watermarking hallucination: factual errors induced or amplified by watermarking even when the required evidence is present in the context and the unwatermarked model can answer correctly. Using a controlled RAG setting, we compare paired unwatermarked and watermarked generations under the same condition. Across representative watermarking methods, including KGW, SWEET, DiPmark, GumbelSoft, Gumbel-Max, and SynthID-style watermarking, we find that watermarking hallucination is widespread: watermarked outputs can remain fluent while introducing factual errors. We attribute this failure mode to two mechanisms: (1) direct token-level bias, which can suppress fact-consistent tokens, and (2) prefix-induced drift, which accumulates through autoregressive decoding and weakens later attention to factual context. Motivated by this analysis, we propose Fact-Preserving Token Intervention (FPTI) and Fact-Preserving Attention Intervention (FPAI), two plug-in interventions that can be integrated into existing watermarking methods to improve factuality. Our experiments show that combining FPTI and FPAI mitigates around 90\% of watermark-induced hallucinations while preserving fluency and comparable decoding efficiency. Overall, this work highlights factuality as a first-class criterion in watermark evaluation, alongside detectability and robustness, and calls for careful factuality validation before deploying watermarks in fact-critical applications.