Mitigating Asymmetric Boundary Encroachment in Continual Learning of Vision-Language Models
Abstract
When learning new concepts sequentially, Vision-Language Models (VLMs) face forgetting that degrades both previously acquired tasks and inherent zero-shot capabilities. This challenge is particularly pronounced in the Cross-domain Task-Agnostic Incremental Learning (X-TAIL) setting, where the absence of explicit task identifiers forces diverse concepts to coexist within a unified and crowded multi-modal space. In this context, we identify an underexplored vulnerability termed Asymmetric Boundary Encroachment (ABE). Even when historical feature spaces remain stable, unconstrained new text features actively encroach upon these established sub-spaces during training. To systematically counteract ABE, we propose a unified framework. First, we introduce Analytic Decision Boundary (ADB), which constructs a geometric defense by enforcing an analytic margin derived from cumulative statistics to secure historical multi-modal boundaries. Furthermore, we integrate Orthogonal Visual Adaptation (OVA) and Analytic Ridge Classifier (ARC) to safely evolve the visual backbone and enhance joint inference. Experiments under X-TAIL setting demonstrate that our framework effectively mitigates ABE and consistently achieves new state-of-the-art performance. Our code is available at https://anonymous.4open.science/r/abe-cl-vlm.