Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
Zidan Wang ⋅ Yaqian Li ⋅ xiaokai zhang ⋅ Kun He ⋅ Kaiwen Long ⋅ Hanpeng Liu
Abstract
CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a resource-efficient route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that the reported forgetting is not caused by insufficient negatives, but by an inappropriate magnitude of the contrastive temperature $\tau$: with $\tau$ set sufficiently small, contrastive loss becomes the strongest single-loss post-training objective. Building on this finding, we propose \textbf{ComCLIP}, a complementary post-training framework that freezes CLIP's text encoder---preserving compatibility with downstream VLMs---and trains the vision encoder under three complementary losses operating in the shared contrastive space: a properly-tempered contrastive loss for image-text alignment, an MSE anchoring loss against the original CLIP for preserving the pretrained feature manifold, and a relational distillation loss from DINOv2 for fine-grained visual knowledge injection. With only one epoch on small data, ComCLIP consistently outperforms prior post-training methods on ViT-B/16 and ViT-L/14 across zero-shot classification, retrieval, linear probing, and MMVP, achieving up to $\mathbf{+7.41}$ points on MMVP. Furthermore, ComCLIP serves as a drop-in vision encoder for LLaVA-7B, improving average performance across $8$ standard VLM benchmarks without any re-alignment of the LLM or projector.
Chat is not available.
Successful Page Load