Learning Robust Representations for Defending White-Box Adversarial Attacks in Continual Learning
Abstract
Continual learning under adversarial perturbations remains underexplored, despite its importance in security-critical deployments where attack patterns evolve over time. Existing continual adversarial defense methods mainly rely on replay or prediction-space regularization, but often fail to preserve robust representations across sequentially arriving attack types. In this paper, we study continual adversarial defense, where a model receives a sequence of tasks induced by different white-box attacks and must acquire robustness to new attacks without forgetting previously learned defenses. We address this challenging learning scenario by proposing the Learning Robust Representation Framework (LRRF), a representation-centric framework that improves robustness through hierarchical alignment at three complementary levels. First, we introduce the Dynamic Representation Matching (DRM) mechanism that aligns clean and adversarial feature distributions within each task to reduce the clean-robustness trade-off. Second, a Cross-Task Invariant Representation (CTIR) mechanism is proposed to regularize the current representation network against an accumulated network from previous tasks, encouraging attack-invariant features to persist over time. Third, a Knowledge Consolidation Optimization (KCO) mechanism is proposed to match class-consistent clean and adversarial features using replayed samples, which further stabilizes category structure across tasks. The empirical results show that the proposed approach achieves state-of-the-art performance.