HalluciText: Mitigating Text Hallucinations in Diffusion-Based Image Restoration
Abstract
Diffusion-based image restoration models produce visually compelling results but systematically hallucinate text---generating characters that appear sharp and legible yet are factually incorrect. As text fidelity emerges as a key requirement for real-world image restoration, this failure mode poses a fundamental barrier to wide deployments of diffusion-based models: a model that replaces degraded-but-faithful text with sharp-but-wrong text misleads users and erodes trust. Recent text-aware models mitigate this by feeding text extracted from intermediate outputs or low-resolution (LR) images into the prompt, but they provide no mechanism to detect or correct the hallucinations that still commonly arise. In this work, we present a systematic study on how to improve text restoration during denoising, demonstrating that standard global conditioning alone is inadequate and that optical character recognition (OCR) confidence scores reliably predict word-level correctness. Grounded in these findings, we propose HalluciText, a training-free framework consisting of two complementary interventions: (i) Confidence-Weighted Classifier-Free Guidance (CW-CFG), which concentrates guidance energy on text regions weighted by OCR confidence to improve text recall; and (ii) VLM-Guided Test-Time Guidance (V-TTG), which leverages a vision-language model (VLM) to identify hallucinated regions and corrects them towards a faithful reference, improving text precision. Both operate solely on the noise prediction or latent representation, making them agnostic to the diffusion backbone and text spotting model. We demonstrate the effectiveness of HalluciText on state-of-the-art (SOTA) text-aware models, achieving up to 5.75 points of improvement in end-to-end text recognition F1 on several datasets without degrading image quality.