Bridging CLIP with DINO: Cross-Modal Information Maximization for Online Test-Time Adaptation
Abstract
Vision-language models such as CLIP enable zero-shot classification from arbitrary class names, but under test-time corruptions, their visual encoder projects distorted features onto incorrect textual anchors with high confidence—a phenomenon we term \textbf{modality bias}. Standard entropy minimization reinforces these overconfident errors, triggering a destructive confirmation-bias loop that collapses the predicted label distribution. We propose \textbf{CDIM} (Cross-Modal Information Maximization), an online test-time adaptation method that breaks this loop at three levels. At the \textit{representation level}, a frozen self-supervised DINO backbone is combined with CLIP via a cross-covariance logit ensemble; by projecting text prototypes into DINO's geometric space—rather than the reverse—the classifier operates free of text bias. At the \textit{optimization level}, a gated information-maximization objective (GT-IM) pairs per-sample sharpening with batch-level diversity, while an adaptive gate excludes unreliable samples from the gradient. At the \textit{prediction level}, an online logit adjustment corrects the residual class-distribution skew. Ultimately, \textbf{CDIM} validates its broad effectiveness by delivering robust, competitive performance across demanding setups, including continual TTA, mixed-domain shifts, and diverse domain generalization benchmarks.