RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent
Abstract
Empirical risk minimization (ERM) is a cornerstone of modern machine learning (ML), supported by advances in optimization theory that ensure efficient solutions with provable algorithmic convergence rates, which measure the speed at which optimization algorithms approach a solution, and statistical learning rates, which characterize how well the solution generalizes to unseen data. Privacy, memory, computational, and communications constraints increasingly necessitate data collection, processing, and storage across network-connected devices. In many applications, these networks operate in decentralized settings where a central server cannot be assumed, requiring decentralized ML algorithms that are both efficient and resilient. Decentralized learning, however, faces significant challenges, including an increased attack surface for adversarial interference during decentralized learning processes. This paper focuses on the man-in-the-middle (MITM) attack, wherein adversaries exploit communication vulnerabilities between devices to inject malicious updates during training, potentially causing models to deviate significantly from their intended ERM solutions. To address this challenge, we propose RESIST (Resilient dEcentralized learning using conSensus gradIent deScenT), an optimization algorithm designed to be robust against adversarially compromised communication links, where transmitted information may be arbitrarily altered before being received. RESIST uses a multistep consensus gradient descent framework with robust-statistics-based screening of neighbor messages. It has a design parameter J that controls the frequency of local gradient computation: each local gradient update is preceded by J-1 communication and robust aggregation rounds. Compared with methods that perform both communication and a local gradient update at every iteration, RESIST with J>2 uses fewer local gradient updates at the same communication budget, which is useful when local gradient computation is the dominant cost. We establish geometric algorithmic convergence guarantees for strongly convex and Polyak–Łojasiewicz ERM problems, and sublinear and finite-horizon guarantees for smooth nonconvex ERM problems. For heterogeneous local objectives, these guarantees quantify neighborhoods for the relevant iterate, objective-value, and stationarity errors. In the strongly convex homogeneous case, convergence becomes exact. We also establish statistical learning-rate guarantees under common-population independent and identically distributed sampling assumptions, including regimes in which the statistical error vanishes as the sample size grows. Experimental results demonstrate the robustness of RESIST across diverse attack strategies, screening methods, and loss functions. The J-ablation experiments further show that larger J can improve convergence and reduce the number of local gradient updates at a fixed communication budget.