On the Instability and Stabilization of Blockwise Muon
Abstract
Muon has become a competitive optimizer for large-scale neural network training through matrix-level orthogonalized updates, but its full-matrix normalization can be expensive. Blockwise Muon is an appealing low-cost approximation: it normalizes matrix shards or small blocks independently, reducing computation and communication overhead while often achieving convergence close to full Muon. However, in large-step-size and large-model regimes, Blockwise Muon can suffer from degraded convergence. We attribute this to cross-block amplification: independent per-block orthogonalizations can collectively amplify updates along certain directions, particularly those governing the effective step size. Motivated by this viewpoint, we accordingly propose \textsc{ACME}: \textbf{A}mplification-\textbf{C}orrected \textbf{M}uon for \textbf{E}ffective step-size recovery. \textsc{ACME} applies a rotational correction so that the directions controlling the effective step size are no longer affected by cross-block amplification. Experiments on language-model pretraining show that, across a range of training settings, \textsc{ACME} achieves convergence nearly indistinguishable from full Muon while retaining the efficiency advantages of Blockwise Muon.