Decentralized $\mu^2$-SGD: Narrowing the Parallelism Gap to Centralized Learning
Ofri Eisen ⋅ Sharon Goldstein ⋅ Ran Elbaz ⋅ Kfir Y. Levy
Abstract
While decentralized learning offers a communication-efficient alternative to centralized distributed training, its scalability is often limited by the number of workers that can be used without degrading statistical efficiency. This limitation is especially pronounced over sparse communication networks, where increasing parallelism can lead to a sharp loss in performance. To address this bottleneck, we introduce *Decentralized $\mu^2$-SGD (DMS)*, a novel decentralized optimization method that significantly extends the parallelism limits of decentralized learning. From a theoretical perspective, within the stochastic convex optimization (SCO) framework, we establish improved bounds on the maximal allowable parallelism, surpassing existing decentralized algorithms. Notably, for a broad class of network topologies, our method matches the parallelism scaling of centralized learning, thereby effectively eliminating the gap between decentralized and centralized optimization. Empirically, we validate our theoretical findings through comprehensive experiments, demonstrating the benefits of our approach.
Chat is not available.
Successful Page Load