Beyond Flat Gossip: Tiered Gossip Learning for Scalable Collaborative AI
Abstract
As collaborative machine learning scales to thousands of edge devices, two failure modes emerge: centralized federated learning creates server bottlenecks and single points of failure, while flat peer-to-peer systems force every node to increase its communication degree as the network grows. We propose Tiered Gossip Learning (TGL), a two-layer push–gossip–pull protocol that decouples data-holding leaf nodes from a decentralized relay backbone. In each round, leaves push local models to a subset of relays, relays gossip among themselves to mix models globally, and leaves pull updates from another random relay subset. This asymmetric design keeps per-leaf communication fixed as the network scales, offloading the mixing burden to a thin layer of higher-capacity relays without relying on any central coordinator for aggregation. Across CIFAR-10, FEMNIST, and AG News, TGL matches or exceeds baseline accuracy with up to 80\% fewer model exchanges. We provide convergence guarantees under standard smoothness, bounded variance, and heterogeneity assumptions, with explicit stage-wise consensus bounds characterizing how relay-layer connectivity governs the global mixing rate.