Anchoring LLM-based Chest X-ray Report Generation via Diffusion Language Planning
Abstract
While Large Language Models (LLMs) have significantly advanced radiology report generation (RRG) with strong generative priors, standard autoregressive decoders still suffer from sequential error accumulation when they must infer all clinical structures directly from images. Diffusion Language Models (DLMs) have recently emerged as a non-autoregressive paradigm with iterative masking and denoising mechanisms. Instead of replacing the mature language decoder with a DLM, we reveal that the key value of DLMs for RRG lies in controllable clinical planning, where the model predicts globally consistent anchor points before free-text realization. To harness this complementary strength, we propose ANCHOR-GEN, the first framework that uses a DLM as an explicit diffusion language planner for RRG. Rather than forcing the autoregressive decoder to infer every critical structure from visual tokens alone, ANCHOR-GEN predicts a structured clinical canvas that serves as clinically grounded anchor points. An autoregressive LLM then leverages these anchors to complete report generation with both fluency and clinical faithfulness. By systematically combining non-autoregressive planning and autoregressive decoding, our unified pipeline is trained with structured denoising, hierarchical coarse-to-fine supervision, and planner-decoder binding objectives. Extensive experiments on the public \textsc{MIMIC-CXR} benchmark demonstrate that ANCHOR-GEN improves clinical faithfulness over prior works. Comprehensive analyses and ablations further show that diffusion planning and causal decoding offer complementary strengths for RRG, and that integrating them in one pipeline provides an effective route toward more faithful medical text generation.