Sign and Magnitude: Two Aspects of Parameter Drift in Continual Learning
Abstract
Studies of loss of plasticity (LoP) in continual learning typically report parameter change through aggregate summaries such as average weight magnitude and dead-unit counts. Such summaries need not reveal the direction of the change or how unevenly the change is distributed across parameters. This paper tracks the sign composition and the full magnitude distribution of weights and biases, layer by layer, as multilayer perceptrons are trained on long sequences of tasks under a variety of activation functions and input preprocessings. Three patterns stand out. The direction of the sign drift depends on the layer, the activation function, and the preprocessing, whereas magnitude growth consistently takes the form of an extending upper tail and is largely insensitive to the preprocessing. The sign composition of the biases is settled almost immediately, and the biases then move considerably further than the weights. These patterns co-occur with dead-unit accumulation and declining accuracy. Together they suggest that parameter drift in continual learning has two separable components, a direction that depends on the layer and the training conditions and a magnitude growth whose form does not, and that this distinction is useful when interpreting aggregate diagnostics of plasticity.