Learning Fine-Grained Vision-Language Alignment from Discriminative Part Descriptions
Abstract
Vision-Language Pre-trained (VLP) models such as CLIP learn strong representations from large-scale image–text pairs and demonstrate impressive zero-shot transfer. However, by learning to align visual and textual features only at a global level, they often suffer from limited interpretability and weak fine-grained perception. This issue stems from pre-training data that overlooks key visual details of object parts, which are often important for distinguishing between subordinate categories (e.g., species of birds or models of cars). Multimodal Large Language Models (MLLMs) built on CLIP-style vision encoders inherit this weakness, limiting both accuracy and trustworthiness. To address this, we propose Part-Aware CLIP (PA-CLIP), a framework that improves fine-grained perception while enhancing interpretability. First, we leverage MLLMs to construct a new dataset,FG-Part, containing about one million part-level image–text pairs that explicitly describe discriminative components (e.g., beak shape or wing patterns). Second, we introduce a part-aware training strategy that encourages explicit grounding of fine-grained textual descriptions to corresponding image regions, strengthening part-level cross-modal alignment. Extensive experiments show that PA-CLIP achieves state-of-the-art results on multiple fine-grained visual recognition benchmarks, validating the benefit of part-level captions for capturing subtle details. Moreover, evaluations on general tasks such as cross-modal retrieval indicate that these improvements do not compromise the model's core generalist capabilities.