ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
Abstract
Contrastive Language-Image Pre-training (CLIP) is fundamentally limited by its 77-token restriction, lack of multilingual capabilities, and coarse-grained semantic representations. While replacing CLIP’s native text encoder with a Large Language Model (LLM)-based embedder offers a promising solution, direct, from-scratch alignment can disrupt the semantic structures learned during pre-training. This often leads to degradation of cross-modal knowledge and particularly impairs recognition capabilities in zero-shot scenarios. In this paper, we propose ProCLIP, a curriculum-learning-inspired progressive alignment framework designed to systematically bridge the LLM embedder and CLIP's visual space, unlocking CLIP's potential for long-text, multilingual, and fine-grained understanding. Our framework operates in two stages: (1) Representation Inheritance, which distills CLIP's original text-space knowledge into an LLM adapter to establish a robust initial vision-language prior, and (2) Contrastive Tuning, which refines the cross-modal alignment with self-distillation regularization on the image encoder to further prevent catastrophic forgetting. To maintain semantic and geometric consistency, we introduce instance-level semantic and global structural alignment constraints throughout both stages. Extensive experiments demonstrate that ProCLIP improves zero-shot classification accuracy by 6.8%--13.5% over existing LLM-augmented baselines and achieves state-of-the-art performance in diverse long-text, multilingual, and fine-grained cross-modal retrieval tasks under comparable settings.