FRESCO: A Novel Consistency Control for Asynchronous Pipeline Parallel Training
Abstract
Asynchronous pipeline parallelism can accelerate distributed model training, but suffers from instability caused by inconsistent gradient computations. We identify a fundamental source of this inconsistency: error signals are coupled to forward-time features, not just parameter states. This means that parameters may evolve as long as they preserve forward-equivalence. Leveraging this insight, we introduce FRESCO, which frames asynchronous updates as a constrained optimization problem: it finds the minimal update modification that acts trivially on the activation subspace. Thus, it serves as a principled generalization beyond the usual synchronous--asynchronous dichotomy: subspace-level protection preserves gradient validity without synchronization stalls, unconstrained-asynchrony instability, or memory overhead that works against pipeline scaling. We demonstrate that FRESCO consistently outperforms representative asynchronous baselines in LLM pre-training, with the gains most pronounced under deep pipelines and high inconsistency levels. Moving beyond the current limitations, FRESCO establishes a new foundation for high-utilization pipeline parallel training, providing a scalable path for large-model development on diverse, resource-constrained systems.