Topological Invariance and Breakdown in Learning Dynamics
Abstract
While mainstream theories of deep learning focus on training with a small learning rate, a growing body of theoretical and empirical work suggests that neural network training dynamics can be qualitatively different at large learning rates. However, a clear and general understanding of the fundamental differences between training at small and large learning rates remains lacking. In this work, we prove that if a group of parameter vectors (neurons) in a neural network model exhibits permutation symmetry, then widely used training algorithms induce a continuous mapping on the set formed by these vectors. Moreover, when the learning rate is below a certain threshold, this mapping becomes a homeomorphism. Our result reveals a critical point of the learning rate: below it, the training dynamics are guaranteed to preserve the topology of the set of neurons, whereas above it, training may introduce simplifications to the underlying neuron manifold. This provides a topological description of a novel implicit regularization effect of small learning rates and offers a potential theoretical lens for empirical phenomena such as loss of plasticity. Notably, our theory is independent of specific network architectures and loss functions, enabling topology to be applied universally to deep learning theory.