RadOmni: Advancing Foundation Model for Non-contrast CT with Omni Radiology Knowledge
Abstract
Non-contrast CT (NCCT) is widely used in clinical practice, yet its limited soft-tissue contrast often renders visual information insufficient for accurate diagnosis. In routine radiology workflows, NCCT is commonly accompanied by contrast-enhanced CT, radiology reports, and pathology reports, which provide complementary knowledge at different levels, from structural and morpho-functional cues to semantic findings and pathological evidence. However, existing pretraining paradigms address this challenge only partially, each introducing complementary but still limited constraints on NCCT representations. To address this limitation, we propose RadOmni, a unified omni-pretraining framework that injects multimodal radiology knowledge into representation learning for non-contrast CT. RadOmni formulates a curriculum consolidation learning strategy, in which knowledge is progressively injected according to optimization difficulty and knowledge level through four stages: masked image modeling on non-contrast CT, cross-phase transfer of anatomical structures and enhanced patterns from contrast-enhanced CT, report-guided semantic learning, and pathology-informed diagnostic enhancement. To mitigate optimization conflicts across heterogeneous objectives and reduce knowledge forgetting, each stage retains and jointly optimizes the objectives from preceding stages, while early encoder layers are selectively frozen to preserve generic representations. To enable such multi-source pretraining, we curate Omni-CT10K, a large-scale CT dataset with more than 10K NCCT scans featuring clinically co-occurring multimodal supervision. Extensive experiments on Omni-CT10K, two in-house benchmarks from different centers, and the public MSD benchmark show that RadOmni consistently achieves strong performance across classification, segmentation, pathological TNM staging, and radiology report generation tasks. Further analyses demonstrate favorable scaling behavior, robust cross-center generalization, and transferable representations.