Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
LIU Hanqing ⋅ Jianjun Cao ⋅ Yuanze Li ⋅ Zijian Zhou
Abstract
Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this phenomenon to intrinsic optimization dynamics, but its triggering mechanism remains unclear. This paper shows that, under commonly used floating-point precision, the Cross-Entropy loss itself can produce a deterministic numerical instability. As training enters a high-confidence stage, the difference between the correct-class logit and the other logits may exceed the absorption-error threshold of float-point number. Then during backpropagation, the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero. This breaks the zero-sum constraint of gradients across classes and introduces a systematic drift in the parameter update of the classifier layer. We prove that this drift forms a positive feedback loop with the feature, causing the global classifier mean and the global feature mean to grow exponentially. We call this mechanism _Numerical Feature Inflation_ ($\mathcal{NFI}$). This mechanism explains the rapid norm growth before a Slingshot spike, the subsequent reappearance of gradients, and the resulting loss spike. We further show that $\mathcal{NFI}$ is not equivalent to an observed loss spike: in more practical tasks, absorption errors may affect only a subset of samples, so spikes do not necessarily occur, while the same zero-sum-breaking mechanism can still drive rapid growth of parameter norms. Our results reinterpret Slingshot as a numerical dynamic of finite-precision training, and provide a testable explanation for abnormal parameter growth and logit divergence in late-stage training.
Chat is not available.
Successful Page Load