Exact Solutions and Saddle-to-Saddle Dynamics in Nonlinear Matrix Factorization
Abstract
Exact spectral analyses of full training trajectories are well developed for linear networks, but nonlinear activations typically destroy the invariant singular-vector structure that makes them tractable. We identify a nonlinear matrix-factorization setting in which this structure survives. For targets with orthogonal columns and elementwise activations, spectral initialization confines gradient flow to an invariant diagonal manifold, reducing training to independent scalar ODEs that are exactly solvable by quadrature. For locally linear activations, small balanced spectral initialization yields sequential mode emergence, with larger target singular values learned earlier, producing trajectories near a sequence of increasing-rank critical points. By contrast, imbalanced spectral initialization synchronizes mode emergence, eliminating the well-separated saddle-to-saddle sequence. We also show that small random initialization exhibits qualitatively similar singular-value dynamics in this model. These results provide an exactly tractable nonlinear setting for studying stage-like feature learning and saddle-to-saddle dynamics.