Temporal Hallucinations in Medical Report Generation
Abstract
Accurately characterizing disease progression in medical report generation is a critical clinical challenge. This paper demonstrates that current evaluation protocols possess a severe blind spot regarding longitudinal reasoning. Using a targeted temporal perturbation dataset, we show that existing metrics consistently fail to penalize crucial temporal errors, assigning misleadingly high scores to trajectory inversions and omissions. To address this, we introduce an LLM-as-a-judge framework to systematically extract and verify comparative clinical claims. Our evaluation of six state-of-the-art models reveals a severe prevalence of temporal hallucinations. Single-image models routinely fabricate longitudinal trajectories, and providing historical imaging proves insufficient, as prior-aware architectures still frequently mischaracterize or omit disease progression. We trace this widespread failure to dataset bias: because ground-truth reports inherently rely on comparative language, models learn to statistically parrot temporal claims to match stylistic priors rather than executing genuine visual reasoning. This work exposes a fundamental weakness in current pipelines and provides a framework for analyzing true longitudinal comprehension.