The Spectral Anatomy of Plasticity Loss
Abstract
Continual learning can make a network progressively worse at fitting new tasks within a fixed training budget. The mechanism behind this loss of plasticity remains unclear. We leverage random matrix theory to follow learned structure as it emerges from the random bulk of the network’s weight matrices. Across an MLP, ResNet-18, and ViT-Tiny, we show that retained function accumulates in persistent directions at the upper spectral edge, forming a fossil record of earlier learning. We further find in ResNet-18 and ViT-Tiny that these directions absorb an increasing share of the gradient for new tasks, while the space of available learning directions progressively contracts. At a representative ViT-Tiny checkpoint, upper-edge directions are enriched in both input covariance and backpropagated error, exposing them to the task on both sides of the layer.