Reinforcement learning enhanced flow matching for reference-based tumor generation on unpaired CT images
Abstract
The success of deep learning in medical image analysis is often hindered by the long-tail distribution of disease cases, where rare or early-stage tumors are critically under-represented. While controllable synthesis offers a potential solution, existing methods suffer from a \textit{specification bottleneck}, as low-dimensional signals like text or binary masks fail to capture complex textures and morphological variations. In this paper, we propose an exemplar-driven synthesis framework that utilizes real tumor samples as high-dimensional templates to enable precise, case-specific generation. To the best of our knowledge, this is the first work to introduce reinforcement learning (RL) into the field of 3D medical image generation, called MedGRPO. We formulate the unpaired tumor synthesis task as a constrained trajectory optimization problem and leverage Flow-GRPO to find an optimal generation path. By designing a multi-objective reward system encompassing texture, shape, and semantic consistency, our model effectively balances pathological fidelity with anatomical coherence without requiring paired supervision. Rigorous evaluations across four major tumor types and the external AbdomenAtlas 2.0 dataset demonstrate that our method produces high-fidelity synthetic data that significantly enhances the performance of downstream diagnostic models.