Biomedical Acquisition-induced Style Shifts as Mixture Shifts: Style-aware Mixture-of-Experts Multimodal Prompt Learning
Abstract
Vision-language models (VLMs), such as CLIP, have shown remarkable transferability in the biomedical domain, and prompt tuning enables efficient adaptation with limited supervision. In practice, biomedical images exhibit pronounced acquisition-induced style shifts, especially across imaging sites and protocols, while cross-modality transfer further introduces compound shifts from different physical acquisition processes. However, existing prompt-tuning methods ignore acquisition-induced style shifts, making the prompts sensitive to source-specific style and limiting the transferability to unseen medical domains. To address this, we formulate acquisition-induced style shifts as mixture shifts, where each domain is viewed as a mixture of latent style components. Under this formulation, the target risk depends on both mixture weights and latent component risks, motivating prompt adaptation that accounts for style-specific variation rather than relying on a single source-specific prompt. In this paper, we propose SaMoE, a Style-aware Mixture-of-Experts Multimodal Prompt Learning framework for adapting VLMs to biomedical domains under acquisition-induced style shifts. To account for latent component risks in prompt adaptation, SaMoE represents multimodal prompts with latent style-conditioned prompt transformations and composes them through image-conditioned routing, enabling each image to induce style-adapted visual prompts. By extracting style cues from intermediate visual representations, the sparse expert composition adapts style-related prompt weights without domain labels and is injected across multiple VLM layers to align with hierarchical visual-language representations. Extensive experiments on 15 medical datasets across 12 modalities and 10 organs demonstrate significant improvements in both accuracy and generalizability over state-of-the-art methods.