FOAM: Factored One-sided Adam-Moment for Practical and Scalable SOAP
Abstract
Second-order optimizers (\eg Shampoo, SOAP) have emerged as compelling alternatives to Adam/AdamW and promise faster per-iteration convergence than AdamW in transformer training. However, they still remain costly in practice: even SOAP still pays substantial optimizer-state memory and eigendecomposition overhead. We ask whether the benefits of SOAP can be retained at nearly the efficiency of AdamW. In this paper, we answer yes with a lightweight yet improved SOAP, Factored One-sided Adam-Moment (FOAM). The path to practical impact is twofold: first, FOAM employs input-only preconditioning, dropping the output-side Kronecker factor to reduce memory and remove one eigendecomposition. A per-row Hessian decomposition identifies this as the right one-sided choice when output curvature is approximately isotropic. Second, FOAM replaces SOAP's dense rotated second moment with a factored Adam-moment estimator based on row-by-column statistics. Under the K-FAC Fisher assumption, this estimator is further proven to be pointwise equivalent to SOAP's update. Our empirical results support the claim that FOAM matches strong SOAP and Muon baselines in validation loss across a diverse scale of large language model (LLM) pre-training, uses optimizer state closer to AdamW, runs within Muon's wall time, and retains robustness across a five-fold learning-rate range. Scaling to 760M parameters, FOAM demonstrates that SOAP's gains need not require SOAP's cost.