Look Before You Leap: Self-Evolving Clinical Reasoning with Psychometric Preference Optimization for Radiology Report Generation
Abstract
Despite the remarkable progress of large Vision-Language Models (VLMs) in Radiology Report Generation (RRG), existing black-box autoregressive generation paradigms lack intermediate clinical logical reasoning. This deficiency renders such models prone to generating linguistically fluent reports with erroneous diagnostic findings, which manifest as plausible hallucinations and thereby undermine clinical reliability. To this end, we propose a novel RRG paradigm termed "Look-Before-You-Leap", aimed at enhancing the clinical logical reasoning capabilities of the model to improve diagnostic accuracy. Specifically, we first construct a self-evolving Radiological Skill (RadSkill) repository to guide the teacher model in distilling high-confidence explicit clinical reasoning, and further internalize this reasoning capability into the implicit diagnostic intuition of the RRG baseline. Building upon this foundation, a novel preference optimization method termed Psychometric Preference Optimization (PSY-PO) is proposed. Grounded in Item Response Theory (IRT) from psychometrics, this method precisely samples clinically misleading dispreferred reports for preference optimization, thereby effectively mitigating plausible hallucinations. Extensive experimental results demonstrate that our proposed method significantly enhances both the generation quality and trustworthiness of radiology reports.