AIR: Rethinking Image-Text Offset Alignment in Multimodal Contrastive Representation Space
Abstract
Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation space (e.g., CLIP or SigLIP). This property is widely utilized in text-guided generative modeling tasks, such as generative model domain adaptation and text-guided Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation spaces such as CLIP and SigLIP. This property plays a central role in text-guided generative modeling tasks, including generative model domain adaptation and text-guided image editing, where textual offsets are used to guide visual transformations. These methods fundamentally rely on the assumption that image and text offsets are well aligned, such that textual transformation directions can provide reliable guidance for corresponding visual transformations. However, the validity of this assumption has not been systematically examined. In this work, we question this foundational assumption by conducting a comprehensive empirical analysis of image-text offset alignment in multimodal contrastive representation space. Our findings reveal not only noticeable offset misalignment but also a meaningful positive correlation between image-text offset misalignment and semantic concept distance across six large datasets and eight contrastive vision-language models. Based on this discovery, we propose Adaptation with Iterative Refinement (AIR), a method that iteratively refines text offsets through anchor sampling and our proposed concept description learning to reduce image-text offset misalignment and improve guidance accuracy. Comprehensive experiments on zero-shot generative model domain adaptation and text-guided image editing, including qualitative, quantitative, and user studies, consistently show that AIR enables state-of-the-art performance in these tasks. Code and additional experiments are available in the supplementary material.