Anatomy-Activated Mixture-of-Experts for 3D Medical Vision-Language Pre-training
Szymon Płotka ⋅ Gizem Mert ⋅ Pedro R. A. S. Bassi ⋅ Wenxuan Li ⋅ Zongwei Zhou ⋅ Anna M Marcinkiewicz ⋅ Wiktoria Romańczyk ⋅ Jarosław B Ćwikła ⋅ Łukasz Struski ⋅ Ewa Szczurek ⋅ Jacek Tabor ⋅ Arkadiusz Sitek
Abstract
Medical Vision-Language Pre-training (VLP) leverages the semantic richness of radiology reports to learn 3D representations, yet current architectures remain fundamentally anatomy-agnostic. While clinical diagnostics rely on organ-specific context, standard encoders apply a uniform, dense parameterization to all volumetric patches, and existing Mixture-of-Experts (MoE) routers gate on isolated local features without explicit access to the organ composition of the scan or crop. We propose Anatomy-Activated Mixture-of-Experts (A$^{2}$MoE), a framework that incorporates anatomical inductive biases within the model's computational path. A$^{2}$MoE introduces two key components: (1) a Histogram Representation Router (HRR) that conditions expert selection on global visual context and a predicted anatomy-presence histogram, ensuring that parameters specialize in specific anatomical domains; and (2) a mask-free Organ-Query Multi-Scale Projector (OQMSP) that utilizes learnable queries to extract organ-level embeddings for fine-grained alignment without requiring manual annotations at inference. By routing at the patch level rather than the token level, A$^{2}$MoE reduces routing complexity by orders of magnitude while satisfying load balancing by construction. Pre-trained on over 60,000 multi-phase CT scans, A$^{2}$MoE achieves state-of-the-art performance across segmentation, classification, and detection benchmarks. Furthermore, we demonstrate that our anatomically-grounded inductive biases support robust generalization to MRI and ultrasound, establishing a scalable foundation for medical representation learning.
Chat is not available.
Successful Page Load