Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
Abstract
Recent empirical studies have shown the training dynamics of deep neural networks mostly occur within low-dimensional subspaces. While this has inspired new research in low-rank training, compression, and adaptation, theoretical justification for these dynamics in nonlinear networks remains limited. To address this gap, this paper analyzes the learning dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD). We demonstrate that the weight dynamics concentrate within invariant low-dimensional subspaces throughout training. Theoretically, we precisely characterize these invariant subspaces for two-layer networks with smooth nonlinear activations, providing insight into their emergence. Experimentally, we validate that this phenomenon extends well beyond our theoretical setting. Leveraging these insights, we empirically show there exists a low-rank MLP parameterization that, when initialized in the appropriate subspaces, nearly matches the performance of their fully-parameterized counterparts on classification tasks.