How Can SignSGD Outperform SGD? A Functional Scaling Law Perspective
Zilin Wang ⋅ Binghui Li ⋅ Jia-Nan Wang ⋅ Lean Wang ⋅ Jinbo Wang ⋅ Lei Wu
Abstract
We study signSGD in linear regression with diagonal features and power-law spectra, a toy but expressive setting where the effect of coordinate-wise sign normalization can be analyzed sharply. Specifically, we establish a functional scaling law (FSL) for signSGD under general learning rate schedules, which decomposes the loss dynamics into two components: learning of the target signal and noise accumulation. Compared with SGD, signSGD learns the signal faster but exhibits slower noise forgetting. This signal-noise tradeoff yields a regime-dependent comparison between signSGD and SGD: signSGD achieves better data-scaling efficiency in hard-task regimes, while they perform similarly in easy-task regimes. The key mechanism is that gradient noise is {\it curvature-aligned}: its directional variance is of the same order as the directional curvature. As a result, sign operation effectively acts as a \(\diag(\bH)^{-1/2}\) preconditioner in the noise-dominated regime, where $\bH$ is the Hessian matrix. Moreover, we demonstrate that the same square-root preconditioning mechanism can arise in RMSprop and Adam, where coordinate-wise normalizers essentially track the coordinate curvatures. Synthetic experiments validate the proposed mechanism and the predicted scaling behavior. Language model experiments further suggest that, beyond the controlled setting, the resulting FSL can serve as a surrogate model for fitting and predicting practical training loss curves.
Chat is not available.
Successful Page Load