Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients?
Zitao Song ⋅ Cedar Site Bai ⋅ Zhe Zhang ⋅ Brian Bullins ⋅ David Gleich
Abstract
Training instabilities such as loss spikes are frequently the result of stochastic gradient noise. Because of rare expressions in language training data, and multiple layer composition, the noise impact is heavy-tailed and survives mini-batch averaging. Gradient clipping is designed to address this where vector-norm clipping ignores matrix structure in weight updates, while spectral normalization (e.g. Muon) respects structure at additional cost. We show an entry-wise \emph{heavy-tailed noise} appears similar to real stochastic gradient noise. Furthermore, through a first-order perturbation analysis, we identify a \emph{localization} property under which a simple entry-wise method will give spectral normalization. Exploiting this, we derive a tractable surrogate for the Bayes-optimal entry-wise estimator under a Gaussian signal prior. We establish $O(\epsilon^{-4})$ convergence guarantee under Cauchy-contaminated noise. Empirically, we find that smooth shrinkage improves Adam on NanoGPT pretraining, saving ${\sim}$8% of training tokens. We further find that applying the entry-wise clipping before spectral normalization yields a ${\sim}$2% token saving on top of Muon.
Chat is not available.
Successful Page Load