Optimizers for Discrete Diffusion Models: A Controlled Benchmark
Abstract
Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present, to our knowledge, the first controlled optimizer benchmark for discrete diffusion models: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8) and uniform diffusion (QM9, and LM1B through the Gaussian duality), each formulation on the task its own paper uses so that every setting has published reference values, plus continuous image diffusion (CelebA-64) as a control. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://anonymous.4open.science/r/diffusion-baselines-B1C7.