DPA: Decentralized Primal Averaging with Quasi-Global Momentum for Highly Heterogeneous Data
Abstract
Decentralized training of deep learning models is widely used to enable data privacy and on-device learning over networks. In realistic scenarios, heterogeneity across clients' local data poses an optimization challenge and can severely degrade test accuracy, especially on sparse communication graphs. Building on recent advances in centralized primal averaging, we propose Decentralized Primal Averaging (DPA), a decentralized learning algorithm that combines primal averaging with a local quasi-global momentum estimate. DPA maintains three coupled sequences for optimization, gradient evaluation, and output averaging, while using consecutive mixed iterates to construct a network-aware momentum direction without transmitting an auxiliary tracking variable. We prove a nonconvex stationarity guarantee for DPA's averaged output model and establish a linear-speedup result up to network- and heterogeneity-dependent residual terms. We also introduce DPA-1G, a one-gossip-per-iteration variant that matches the communication budget of DSGD-style baselines. Experiments following the GUT decentralized image-classification benchmark show that DPA outperforms DSGD, gradient tracking, QG-DSGDm, momentum tracking, and QG-GUTm across datasets, architectures, topologies, and heterogeneity levels, with the largest gains under high heterogeneity. DPA-1G matches or exceeds the one-gossip baselines at the same per-iteration communication budget.