Generalization Without Compression Penalty: A Stability Analysis of Error Feedback
Yifei Liang ⋅ Peng Wang ⋅ Yan Sun ⋅ Yingqi Liu ⋅ Xiaoyan Wang ⋅ XIAOCHUN CAO ⋅ Li Shen
Abstract
Error Feedback (EF) has become a core mechanism for aggressive gradient compression in communication-efficient distributed training, and is now used in geo-distributed frameworks such as DiLoCo. While existing theory mainly explains its optimization convergence, the generalization behavior of EF remains poorly understood, especially under nonlinear biased compression and distributed local updates. We present the first generalization analysis of EF for both single-node EF-SGD and DiLoCo-EF through an optimization-driven on-average model stability framework, without imposing bounded-gradient assumptions. Technically, we develop a buffer-aware Lyapunov argument that tracks the coupled dynamics of the model and error buffer, while co-coercivity absorbs same-sample gradient-difference energy into the optimization trajectory. For convex losses, our bounds show that compression-induced stability terms are transient, decaying as $\mathcal{O}(\sqrt{b_\alpha}\,T^{-3/4}+b_\alpha T^{-1})$, and therefore do not create a non-vanishing generalization floor beyond the dominant $\mathcal{O}(T^{-1/2})$ optimization term. In the DiLoCo setting, worker averaging further suppresses compression perturbations, and a local stepsize $\eta_l=\Theta(1/(H\sqrt{T}))$ controls the drift from $H$ local steps, yielding the leading $\mathcal{O}(1/\sqrt{MHT})$ excess-risk scaling. Experiments on convex benchmarks corroborate the theory, showing that Top-$k$, Random-$k$, and quantization with matched contraction factor $\alpha$ exhibit nearly indistinguishable generalization trajectories.
Chat is not available.
Successful Page Load