A Mahalanobis Margin \texorpdfstring{$\gamma_{\min}$}{gamma\_min} Bound on Task Confusion in Pretrained Class-Incremental Learning: From Infeasibility to Exponential Attenuation
Milad Khademi Nori ⋅ Guanghui Wang
Abstract
Class-incremental learning (class-IL) faces a structural obstacle: discriminative sequential training optimises only the diagonal blocks of a pairwise loss matrix, leaving the inter-task blocks---task confusion (TC)---unminimised. A recent Infeasibility Theorem makes this precise, proving that no discriminative class-IL learner can attain the joint-training optimum even when catastrophic forgetting (CF) is perfectly addressed. Yet the state of the art consists almost entirely of discriminative methods built on pretrained foundation models that approach joint-training performance. We explain this regime by showing that pretraining \emph{exponentially attenuates} the obstacle. For any pretrained backbone $\phi$, the TC gap of any CF-optimal prototype-based discriminative learner---a class that captures the test-time classifier of SLDA, RanPAC, FeCAM, and the prototype branches of EASE and InfLoRA---is bounded by $\exp(-\gamma_{\min}(\phi)/8)$, where $\gamma_{\min}(\phi)$ is the worst-case Mahalanobis margin between task centroids in feature space: a single, computable scalar that determines how much TC remains. Combined with a finite-sample convergence analysis under heavy-tailed gradient noise and a representation drift bound governing backbone fine-tuning, this yields a diagnostic decomposition of the class-IL excess risk into $E_{\mathrm{CF}} + E_{\mathrm{TC}}$, with disjoint direct dependencies on optimisation budget and backbone quality. Across 12 (backbone, dataset) configurations spanning ViT-IN21K, DINOv2, and CLIP plus four random-backbone anchor points, the pooled empirical slope of $\log(\mathrm{TC\ gap})$ versus $\widehat\gamma_{\min}$ on softmax cross-entropy heads is $-0.119$ (bootstrap 95\% CI $[-0.134, -0.105]$), consistent with the predicted $-1/8$ and excluding the conservative fall-back $-1/16$. The slope fit extends the rate beyond the formal LDA/QDA scope. $E_{\mathrm{CF}}$ scales as $K^{-(\alpha-1)/\alpha}$ as predicted, with the rate steepening by more than $3\times$ under gradient clipping; representation drift scales linearly in the backbone learning rate ($r = +0.998$). A theory-derived training procedure---whose schedule is determined by $\widehat\gamma_{\min}$, the tail index $\widehat\alpha$, and warm-up estimates of smoothness and curvature---is competitive with the strongest pretrained-CIL baselines across the four standard benchmarks.
Chat is not available.
Successful Page Load