Why Heavy-Tailed Weights Predict Model Quality
Abstract
A solid understanding of the predictive success of deep neural networks (DNNs) remains elusive. For example, while many data-dependent metrics exist that seek to predict model quality of DNNs, these metrics may be expensive to compute for large datasets and/or may be impossible to compute when training datasets are not released. A practical and effective predictor of model quality involves analyzing DNN weight matrices, as it is known that a heavy-tailed spectral distribution often correlates with strong model quality. As such metrics lack a rigorous statistical derivation, in this work we aim to understand why they perform well, by employing a probabilistic framework to derive the marginal likelihood for the trained weights of a DNN. Across large-scale convolutional DNNs and LLMs, we find that this marginal likelihood predicts model quality. To understand the role heavy-tailed spectra play in predictive performance, we derive the limiting log-marginal likelihood for spectra that follow the heavy-tailed High-Temperature Marchenko-Pastur (HTMP) distribution. We show that this limiting quantity is convex in a heaviness parameter, and we derive an optimal heaviness that empirically predicts over-fitting for large-scale DNNs during training.