An Equivariance Principle for Optimizer Design: Symmetry-Compatible Updates for Embeddings, LM Heads, and MoE Routers
Abstract
Most modern deep neural networks are trained with Adam and related coordinate-wise adaptive optimizers, which are efficient but do not respect the natural matrix symmetries, geometry, and spectral structure of neural network parameters. We develop a symmetry-based principle for matrix-gradient optimization, showing that orthogonally equivariant optimizers, especially spectral optimizers such as stochastic spectral descent, Muon, Scion, and polar gradient methods, are better aligned with matrix layers than coordinate-wise updates. We show that coordinate-wise adaptivity breaks orthogonal equivariance, discards gradient spectral structure, and can be viewed as applying spectral updates to a pathological diagonal lifting of the parameter space. Motivated by permutation and orthogonal symmetries, we derive symmetry-compatible optimizers for different layer types, including linear layers, embedding and LM head matrices, and MoE routers. This yields one-sided spectral optimizers, row-norm optimizers, and hybrid row-norm/one-sided-spectral variants. Pre-training experiments on dense and sparse MoE language models, including Qwen3-0.6B-style, Gemma 3 1B-style, OLMoE-1B-7B-style, and downsized gpt-oss models, show that replacing AdamW on large vocabulary-indexed matrices and router matrices with symmetry-compatible optimizers consistently improves final validation loss and sometimes training stability.