DeltaMomentum: A Key-Value based Anisotropic Momentum Update via Delta Rule
Euijin Hong ⋅ Guannan Qu
Abstract
The first-moment buffer of most modern optimizers is an exponential moving average (EMA) of stochastic gradients with a single scalar decay, with a unified forgetting horizon along every direction in parameter space. Yet the per-sample gradient of any linear layer factorizes as a rank-1 outer product $g_t = \delta_t x_t^\top$, exposing the input activation as a natural $\textit{key}$ and the output-side error as a natural $\textit{value}$, which is the structure that EMA's flat matrix average discards. We propose $\textbf{DeltaMomentum}$, which interprets the buffer as an online linear associative memory of these key--value pairs and updates it via the classical delta rule. The resulting update is $\textit{anisotropic memory transport}$: directions queried frequently are forgotten quickly, directions queried rarely are preserved on long horizons, yielding an automatic, data-driven schedule matched to input statistics. We prove the buffer's expected fixed point is a Tikhonov-regularized Wiener predictor of the population gradient, which is equivalent to implicit input-side natural-gradient preconditioning at no covariance-inversion cost, and that the same mechanism strictly accelerates expected tracking-error contraction along every positive-density input direction under non-stationarity and reduces the per-direction iteration determinant in linear regression. DeltaMomentum is a drop-in replacement for the EMA accumulator in any base optimizer; we instantiate it as DeltaSGD and DeltaAdamW. A $\mu$P derivation shows the delta coefficient is width-invariant, enabling zero-shot hyperparameter transfer; we verify this with a coordinate check across five widths. The per-block asymptotic FLOP overhead is $\approx 15.74$% at zero persistent-memory cost (down to $7.3$% and $11.0$% realized at 67M/370M). On FineWeb-Edu, DeltaAdamW reaches AdamW's terminal validation loss in up to $31.35$% fewer steps at 67M and $19.28$% fewer steps at 370M. On CIFAR-10, DeltaSGD shows the same pattern against SGD-momentum, confirming the gain is not specific to Adam. Mechanistic diagnostics confirm that improvements in gradient-estimator quality, function-space prediction error, and input-feature conditioning are operative throughout training.
Chat is not available.
Successful Page Load