MedIGen: Reliable Medical Illustration Generation via Interleaved Introspective Reasoning
Abstract
Medical illustration generation is not merely domain-specific text-to-image (T2I) generation, but a high-constraint visual reasoning problem where visually plausible images can be invalid due to anatomical and structural errors. Existing T2I models largely rely on one-pass generation and lack an internal mechanism to inspect and correct such biomedical inconsistencies. We introduce MedIGen, a unified medical illustration generation model that transforms one-pass synthesis into an interleaved introspection-aware refinement process, coupling generation with self-reflection and re-generation through “generate–reflect–refine” cycles. MedIGen is trained with three progressive stages: (1) large-scale medical illustration pretraining on 1.21M samples, (2) mixture training over generation, reflection, and refinement sub-skills, and (3) reinforcement learning with dual-level rewards that optimize both reasoning validity and final visual correctness. We further present IlluGenBench, an expert-aligned, rubric-driven benchmark with 296 diverse tasks and 9,015 criteria, evaluating scientific accuracy, structural correctness, and semantic alignment beyond coarse visual plausibility. Experiments show that MedIGen substantially outperforms strong open-source unified and reasoning-based baselines, establishing a new open-source frontier. All resources will be open-sourced to facilitate future research.