Decoupled Mode Connectivity for Base-to-Novel Generalization in Vision-Language Models
Abstract
Prompt learning adapts vision-language models such as CLIP by optimizing a small set of continuous context vectors. The cross-entropy objective drives the learned prompt toward base-class specialization, while generalization to unseen classes benefits from staying close to the zero-shot feature space. Because both objectives act on the same parameters, their gradients oppose each other at convergence, and single-prompt losses are confined to a fixed empirical base-novel trade-off curve across a wide range of loss designs. We propose Decoupled Mode Connectivity (DMC), which resolves this conflict by assigning each objective to a dedicated prompt. A linear mode connectivity (LMC) corridor in text-feature space enforces low classification loss across interpolated classifiers between the two endpoints. The Visual Anchor regularizer preserves CLIP's pretrained class-similarity structure during specialization; equivalently, it minimizes the KL divergence between the learned and pretrained class-similarity distributions. We introduce class-permutation invariance (CPI) as a necessary condition for regularizer transfer across the base-novel boundary, and prove via Fano's inequality that any CPI violation lower-bounds the drop in novel accuracy by a term proportional to the mutual information between the prompt and base-class labels. DMC improves base and novel accuracy on the majority of dataset-baseline combinations across 11 datasets and three prompt-learning baselines~(CoOp, KgCoOp, MMA), shifting the Pareto frontier rather than trading along it.