SGD at the Edge of Stability: The Stochastic Sharpness Gap
Fangshuo Liao ⋅ Afroditi Kolomvaki ⋅ Anastasios Kyrillidis
Abstract
Edge of Stability (EoS) refers to the phenomenon where full-batch gradient descent (GD) training of neural networks with step size $\eta$ pushes the largest eigenvalue of the Hessian, i.e. the sharpness, to $2/\eta$ and hovers there. \citet{damian2023selfstab} explains the hovering behavior of the sharpness by *self-stabilization*, a mechanism driven by third-order structure of the loss, and shows that GD implicitly follows projected gradient descent (PGD) on the set where the sharpness is below $2/\eta$. For mini-batch stochastic gradient descent (SGD), the sharpness stabilizes *below* $2/\eta$, with the gap widening as the batch size decreases. However, no theoretical explanation exists for this suppression. In this paper, we introduce *stochastic self-stabilization* that extends the self-stabilization framework to SGD. Our key insight is that gradient noise injects variance into the oscillatory dynamics along the top Hessian eigenvector, strengthening the sharpness-reducing force and shifting the equilibrium below $2/\eta$. Following the approach of \citet{damian2023selfstab}, we define *stochastic predicted dynamics* that tracks the deviation of SGD from the PGD trajectory, and prove a stochastic coupling theorem that relates the SGD sharpness and loss to those of PGD. Based on our predicted dynamic, we derive a closed-form equilibrium sharpness gap that scales with the variance of the gradient noise projected onto the top eigenvector of the Hessian. This formula predicts that smaller batch sizes yield flatter solutions, and recovers GD when the batch equals the full dataset.
Chat is not available.
Successful Page Load