Revisiting the Adam–SGD Gap Beyond Single Factors
Abstract
Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data properties, architecture design, and optimization dynamics. Yet these explanations are often studied in isolation, leaving their relative importance unclear. In this work, we revisit these hypotheses through a controlled empirical study across vision, language, genomics, and graph tasks, spanning modern and classical architectures, and carefully designed training setups. Our results suggest that no single factor consistently explains the Adam--SGD gap. For instance, the Adam advantage can (1) persist under a uniform vocabulary distribution yet nearly disappear under a heavy-tailed one; (2) reverse in favor of SGD in softmax-attention models; and (3) become larger when ReLU is replaced by GeLU nonlinearity. Instead, the gap emerges from interactions between data and architectural properties. Yet, we observe a consistent pattern across our settings: a crossover batch size at which the relative advantage shifts from SGD to Adam as batch size increases. This perspective helps reconcile several existing hypotheses while offering practical insights across domains.