Discovering and Guiding Clinical Capability Transfer in Medical Vision-Language Models
Ananya Shukla
Abstract
A central but poorly understood question in medical vision-language learning is not simply whether a model can acquire diverse clinical capabilities, but how learning one capability changes its ability to acquire another. Medical VLMs are increasingly trained and evaluated across classification, retrieval, grounding, reasoning, and generation, yet these capabilities are largely treated as independent objectives. Therefore, we lack a principled account of whether clinical supervision creates reusable capability, remains task-specific, or induces interference when transferred across tasks. We investigate this problem as clinical capability transfer, using 3D CT as a controlled testbed where supervision spans anatomical structure, finding localization, pathology recognition, image-text alignment, grounded reasoning, and report-level synthesis. We introduce a framework that treats each task $T_i$ as a controlled intervention on a shared CT-VLM. Task Intervention performs parameter-efficient adaptation to a source task, after which its effect on a held-out target is measured relative to the original model, yielding a directed transfer matrix that explicitly captures positive, negative, and asymmetric transfer. Transfer Attribution characterizes each task by the clinical evidence its supervision requires and models transfer as a function of source-target supervision structure, distinguishing systematic capability transfer from superficial task similarity or source-task strength. Transfer-Guided Adaptation then uses the estimated transfer structure to select auxiliary supervision that is most beneficial for a given target capability, testing whether this improves adaptation when only limited target supervision is available. Using CT-RATE and RadGenome-ChestCT with contemporary 3D CT-VLM backbones, we evaluate transfer across global, local, diagnostic, and image-language capabilities, including transfer asymmetry, negative-transfer pathways, cross-task generalization, and robustness across model backbones. The framework is designed to distinguish three scientifically meaningful regimes: transferable capabilities that provide reusable clinical knowledge, task-specific capabilities that require direct supervision, and competing capabilities whose objectives interfere under shared adaptation. We hypothesize that clinical supervision induces a structured and directional transfer landscape rather than a uniformly shared capability space, and that this structure can predict which auxiliary tasks are most valuable for a given downstream capability. By connecting clinical supervision to measurable capability transfer and, ultimately, training decisions, this work establishes a principled framework for understanding and exploiting how medical VLMs acquire interdependent clinical capabilities under constrained supervision.
Chat is not available.
Successful Page Load