Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
Arseniy Andreyev ⋅ Pierfrancesco Beneventano
Abstract
Recent findings by \citet{cohen_gradient_2021} demonstrate that when training neural networks with full-batch gradient descent at step size $\eta$, the largest eigenvalue $\lambda_{\max}$ of the full-batch Hessian consistently stabilizes around $2/\eta$. These results have significant implications for convergence and generalization. This stabilization, however, does not occur for mini-batch stochastic gradient descent (SGD), so the implications above do not directly transfer. We show that SGD trains in a different regime we term Edge of Stochastic Stability (\textsc{EoSS}). In this regime, what stabilizes at $2/\eta$ is \emph{Batch Sharpness}: the expected directional curvature of mini-batch Hessians along their corresponding mini-batch gradients. As a consequence, $\lambda_{\max}$---which is generally smaller than \emph{Batch Sharpness}---is suppressed, aligning with the long-standing empirical observation that smaller batches and larger step sizes favor flatter minima. We further discuss implications for mathematical modeling of SGD trajectories.
Chat is not available.
Successful Page Load