GRAPE: Knowledge Graph–Derived Reasoning for Multimodal PE Report Generation
Abstract
Pulmonary Embolism (PE) is a life-threatening condition that requires rapid and accurate diagnosis from Computed Tomography Pulmonary Angiography (CTPA) scans. Although recent multimodal large language models (LLMs) have shown promising capabilities in radiology report generation, they frequently suffer from clinical hallucinations and lack interpretable diagnostic reasoning grounded in standardized medical knowledge. Existing approaches primarily rely on latent statistical correlations and do not explicitly model the intermediate reasoning process followed in clinical diagnosis. To address these limitations, we propose a knowledge-centric multimodal framework for pulmonary embolism report generation using the INSPECT dataset. The proposed framework introduces a graph-derived reasoning supervision strategy in which imaging-derived findings are mapped to UMLS concepts and used to retrieve multi-hop reasoning trajectories from a medical knowledge graph through constrained graph traversal. These reasoning paths are transformed into structured chain-of-thought supervision signals that explicitly connect imaging findings, anatomical structures, and pathological conditions, enabling the model to learn clinically grounded diagnostic reasoning during training. The generated reasoning chains subsequently guide and constrain the language model toward interpretable and diagnostically consistent radiology report generation. Experimental results on approximately 23,000 CTPA studies demonstrate that the proposed framework outperforms existing multimodal and reasoning-based baselines, achieving a BLEU-4 score of 0.601, ROUGE-L score of 0.772, and BERTScore F1 of 0.963. Human and LLM-based evaluations further indicate that the framework produces clinically meaningful reasoning and factually consistent radiology reports.