EIPO: Efficient Inter-Step Parallel Optimization
Jianrong Lu ⋅ Zhuoya Gu ⋅ Zhiyu Zhu ⋅ Hui LIU ⋅ Junhui Hou
Abstract
The wall-clock cost of training large neural networks is dominated by the long horizon of sequential optimizer iterations. We recast the gradient-descent (GD) trajectory of modern optimizer (e.g., SGD and Adam) as the unique root of a nonlinear operator, enabling a parallel root-finding approach that updates multiple GD steps simultaneously. We then propose \textbf{E}fficient \textbf{I}nter-step \textbf{P}arallel \textbf{O}ptimization (\textbf{EIPO}) for accelerating a wide range of optimizers. EIPO contributes three pieces: (i) a Preconditioned Nonlinear-Equations (PNE) formulation that recovers the exact sequential GD trajectory; (ii) Memory-Efficient Adaptive Anderson Acceleration, which extracts the multisecant geometry directly from the optimizer's intrinsic momentum buffers and therefore avoids the memory overhead of classical Anderson Acceleration. (iii) We prove that EIPO converges to the optimizer's GD trajectory within fewer iterations than its sequential counterpart. (iv) Across language modeling on WikiText-2 (GPT-2, Llama-3.2-1B, Qwen-1.5-4B/2.5-3B, Gemma-2B, GPT-J-\textbf{6B}), image classification on CIFAR-10 (ResNet50, ViT, CNN), and diffusion training on LSUN Church (UNet/DDPM/DDIM), EIPO reduces optimizer iterations by up to $\textbf{21}\times$ and wall-clock time by up to $\textbf{4.6}\times$ while matching baseline perplexity, accuracy, and FID, all under comparable GPU memory and token consumption. Source code is available at: \url{https://anonymous.4open.science/r/EIPO-44BC}.
Chat is not available.
Successful Page Load