ASSET: Acquisition-Sensitive Subspace Estimation for Test-Time Adaptation of Medical VLMs
Abstract
Medical vision-language models are often adapted under high-quality acquisition conditions. At target deployment, the same model may face lower-quality scanners, protocols, reconstructions, or artifacts. These shifts change the visual evidence, while the intended clinical question remains unchanged. As a result, a source-adapted model can fail under this acquisition shift. Methods are proposed to solve this problem. Medical acquisition-shift correction mainly targets image-only predictors, leaving the multimodal model underexplored. Robust training depends on anticipated degradations, and generic test-time tuning can easily overfit the small target training set, leaving unknown acquisition shifts insufficiently corrected. In this paper, we hypothesize that acquisition-like augmentations of target examples can expose parameter directions responsible for acquisition-shift failures. Based on it, we propose \textsc{ASSET}, an \emph{Acquisition-Sensitive Subspace Estimation} method for test-time adaptation. ASSET estimates an acquisition-sensitive parameter subspace from prediction and gradient changes under these augmentations. It then adapts the model within this subspace through iterative estimate-and-update rounds. Comprehensive experiments show the advantages of ASSET. For example, ASSET consistently improves target accuracy under acquisition-quality shift across datasets and backbones by 2.62\% on average, while preserving performance when no shift is present. Under the strongest acquisition-quality shift, ASSET even improves target accuracy by 4.12\%, showing that the correction generalizes to severe shifts.