Complementary Cache Guidance with Gradient Disentanglement for Continuous Test-Time Adaptation
Abstract
Continuous test-time adaptation (CTTA) is crucial for deploying vision-language models (VLMs) in real-world environments where data distributions shift over time. Existing approaches typically utilize confidence-driven pseudo-labeling to enhance the performance on the target domain while retraining VLMs using historical information. Despite the progress, their performance is far from satisfactory due to error accumulation from noisy pseudo-labels and optimization interference between newly acquired and historical knowledge. Towards this end, we propose a novel approach named Complementary Cache Guidance with Gradient Disentanglement (CURE) for CTTA of VLMs. The core of our CURE is to balance instant knowledge acquisition and historical knowledge consolidation using both complementary cache systems and parameter space disentanglement. In particular, our CURE first constructs affinity structures among diverse text prompts, which guide majority voting to improve the quality of pseudo-labeling. More importantly, our CURE introduces a complementary cache system consisting of a short-term cache and a long-term cache with sample quality monitoring. The short-term cache stores recent reliable samples with high entropy, while the long-term cache preserves representative samples with high gradient consistency across iterations. Then, we extract prototypes from both caches for cross-modal alignment, enabling the model to acquire new knowledge while mitigating old knowledge forgetting. To further reduce the interference when learning from two caches, we perform low-rank decomposition of the gradient space, which facilitates historical knowledge consolidation in the null space of short-term signals. Extensive experiments on several benchmarks demonstrate the superiority of CURE over state-of-the-art baselines. Our source codes are available at https://anonymous.4open.science/r/CURE-B486.