Text-side fragility in contrastive vision-language retrieval
Yanyu Lei
Abstract
Contrastive vision-language models (VLMs) are deployed in a growing range of retrieval and grounding systems, making their robustness to input corruption a critical property. In this paper, we apply controlled corruptions to textual and image inputs for five contrastive VLMs spanning three Vision Transformer architectures, four training corpora, and two contrastive loss functions. We aim to measure how image and text representations and retrieval performance respond as input damage increases. Previous cross-modal robustness work often applied ordinal severity scales to corruptions independently, even though the same severity step on the image side may not represent equivalent damage on the text side or vice versa. Our experiment addresses this measurement gap using within-modality damage percentile binning, which ranks each side's corruptions on its own damage scale and compares only at matched ranks. We evaluate all models on MS-COCO retrieval with $n = 1{,}000$ pairs over three seeds, six corruption types, and five severity levels. Our results indicate that image-to-text (i2t) Recall@1 at the highest damage quintile is 4.6$\times$ to 5.8$\times$ higher than text-to-image (t2i) Recall@1 at the matched quintile. Two mechanistic probes are used to localize part of the asymmetry to the single-token text-pooling bottleneck and to text damage that destroys lexical content rather than word order. These results suggest that downstream applications retrieving images from text are notably more fragile than the reverse, and that future improvements to cross-modal robustness require architectural attention to text readout.
Chat is not available.
Successful Page Load