AcceleGrad#: Adaptive Geometry-Aware Acceleration
Hanka Goralija ⋅ Francesco Tonin ⋅ Kimon Antonakopoulos ⋅ Alp Yurtsever ⋅ Volkan Cevher
Abstract
Modern neural-network optimizers occupy a three-way design space: online adaptivity, architecture-aware geometry, and acceleration. We introduce AcceleGrad#, an accelerated linear-coupling template that couples an adaptive anchor with a norm-induced sharp, mirror, or linear-minimization geometry branch. The scalar-anchor specialization recovers the accelerated $O(1/K^2)$ rate for smooth convex objectives, while the diagonal-AdaGrad clipped-LMO training variant admits local $\mu$-Kurdyka-Lojasiewicz certificates: a generic conservative $O(K^{-1/4})$ rate and, in a SCION cross-entropy regime, a $O(K^{-1/3})$ last-iterate rate to a self-bounded mini-batch noise level. On image classification and language modeling, the training variant improves over Euclidean AcceleGrad and remains competitive with recent non-Euclidean optimizers. Using stochastic gradients directly in spectral oracles also reduces polar-computation time and optimizer overhead in practice.
Chat is not available.
Successful Page Load