EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
Abstract
Lookahead-based acceleration methods, such as Nesterov’s momentum, are widely used in optimization, but they often become unreliable in deep learning training due to stochastic gradient noise and non-convex loss landscapes. Standard lookahead relies on short-horizon update signals (e.g., differences between consecutive iterates), which are inherently noisy and can lead to unstable extrapolation directions. This work revisits Nesterov's acceleration from a trajectory perspective and argues that effective acceleration in deep learning should follow the low-frequency trend of optimization trajectories rather than extrapolating noisy one-step updates. Leveraging on this insight, we propose EMA-Nesterov, a simple modification that replaces the standard Nesterov's lookahead direction with an exponential moving average of parameter updates. This yields a stabilized lookahead direction that approximates long-horizon trajectory trends while retaining a lightweight single-loop implementation. We show that EMA-Nesterov retains the same theoretical accelerated convergence rate in convex problems. Furthermore, we show that EMA-Nesterov consistently improves optimization performance across a range of optimizers, including Adam, SOAP, and Muon, in deep learning tasks such as language model pre-training. Compared to prior lookahead methods, EMA-Nesterov achieves better performance by avoiding the instability of short-horizon lookahead and inefficiency of multi-step lookahead.