Beyond Hard Negatives: Grounded Positive Supervision for Compositional CLIP
Abstract
CLIP-style vision-language models transfer broadly but remain brittle for compositional reasoning, often collapsing object-attribute binding and inter-object relations. We attribute this gap to limited compositional diversity in web-scale data and contrastive objectives that prioritize global alignment over grounded structure. Many fixes rely on caption-edited hard negatives, which can bias learning toward specific perturbation templates and trade off against general alignment. We propose GPS-CLIP, which improves compositionality by adding multi-granular positive constraints to standard contrastive learning: IoU-weighted multi-positive region-span alignment and relation-triplet alignment in the shared embedding space. GPS-CLIP uses existing grounded annotations during fine-tuning, keeps the dual-encoder architecture unchanged, and requires only images and text at inference. Across SugarCrepe, What'sUp, COLA, and Winoground, GPS-CLIP achieves state-of-the-art compositional performance while improving cross-modal retrieval and zero-shot recognition. Beyond compositionality, GPS-CLIP (i) achieves the strongest MMVP-VLM results for fine-grained visual understanding and (ii) yields a more structured embedding space with improved object/attribute separability on UT-Zappos geometric analysis compared to hard-negative variants. Overall, GPS-CLIP offers a streamlined way to improve VLM compositional robustness.