A Closed Form Depth-Width Law for Neural Collapse in Deep Linear Networks
Pankhil Gawade ⋅ Kirsten Fischer ⋅ Guy Wolf ⋅ Ewa Szczurek
Abstract
How depth and width of neural networks shape learned representations is usually studied only after training. We show that, in a tractable feature-learning setting, this dependence can instead be obtained in closed form. Using neural collapse as a probe of representation geometry, we express NC1-NC3 directly in terms of the learned feature kernel, extending an identity known for NC1 to the angular part of NC2 and to NC3 under a ridge readout. For deep linear networks with exchangeable class structure, the resulting kernel equations yield a geometric law, exact within the theory, for the layerwise evolution of within-class variability. At large width, the accumulated compression becomes a function of the depth-to-width ratio $L/N$, with its rate determined self-consistently by the data, prior and architecture. Numerical experiments support the derived dependence on depth, width and signal-to-noise ratio. The result is a representation-level scaling law derived from feature-learning theory rather than fitted after training.
Chat is not available.
Successful Page Load