Class-Mixed Diffusion Augmentation for Shortcut-Breaking in Continual Learning
Abstract
Zero-shot pre-trained diffusion models can provide a training-free source of synthetic image generation for Exemplar-Free Class Incremental Learning (EFCIL). However, same-class generation is limited by the distribution mismatch with real data. Parameter-allocation-based EFCIL can minimize the forgetting in model parameters, yet the performance degrades because task-id prediction fails when task-specific feature extractors overcompress and converge to shortcut-dominant representations. We propose \emph{Class-Mixed Diffusion Augmentation} (CMDA) that uses diffusion models not only as past-task data generators but also as controlled generators that challenge the shortcut reliance. We sample candidate tokens from the vocabulary and select \emph{shortcut-breaking} tokens via classifier-guided filtering, then treat the selected tokens as auxiliary classes during training. We provide theoretical analysis linking shortcut sharing to an intrinsic lower bound on task-id error and show that minimum-CE selection recovers shortcut-breaking tokens. Experiments on Split-CIFAR100 and Split-ImageNet show consistent gains over EFCIL baselines, with modest overhead over vanilla diffusion augmentation and far lower cost than diffusion fine-tuning or gradient-guided sampling.