Evaluating Report-Derived Supervision Strategies for Robust Medical Vision-Language Models Under Temporal and Institutional Shift
Abstract
Medical vision-language models increasingly rely on report-derived supervision, but its effect on robustness is poorly understood. We compare seven structured-label and text-based supervision strategies for X-ray hardware recognition across internal, temporal, and external evaluation. Phrase-level contrastive supervision was consistently most robust on the primary task, particularly under temporal and institutional shift. However, the best supervision strategy depended on the downstream endpoint, with different approaches favored for fine-grained classification and retrieval. Pooled calibration error remained low overall but degraded under shift, and fixed high-sensitivity operating points transferred poorly across institutions. These results show that supervision design affects both robustness and what information a representation preserves, motivating evaluation across distribution shifts and downstream tasks.