FAUST: Federated Asynchronous Update with Staggered Timescales for Low-Communication Foundation Model Training
Yunlong Tan ⋅ Mingqiao Mo ⋅ Hao Zhang ⋅ Shengxing Qin
Abstract
Large language model training with sequence-parallel data-parallel (SP-DP) is bottlenecked by the interplay between intra-step activation redistribution and inter-step gradient synchronization. Existing infrequent-communication methods like FedAvg keep states local or reset them lack convergence guarantees and are unstable under long-context sequence-parallel regimes where each device holds only a partial activation shard; Local Adam synchronizes all states jointly and is convergent but triples the collective payload, and its uniform synchronization period cannot distinguish the semantically mandatory intra-step All-to-All from the temporally deferrable inter-step AllReduce. We propose FAUST (Federated Asynchronous Update with Staggered Timescales), a compiler-orchestrated training system that composes compile-time sequence-parallel graph rewriting with tiered runtime synchronization, assigning independent averaging periods to parameters ($K_x$), first moments ($K_u = 3K_x$), and second moments ($K_v = 6K_x$) according to their impulse-response half-lives, while preserving convergence under the orthogonal collective constraint that intra-step All-to-All must retire before inter-step AllReduce begins. Our analysis shows that while first-moment drift dominates the convergence rate in-distribution, high-fidelity convergence guarantees require at least periodic synchronization of the second moment. Experiments on language models up to 1.3B parameters show that FAUST incurs $303{\times}$ less wall-clock communication overhead than standard DDP and $1.78{\times}$ less than the previous state-of-the-art DES-LOC, while achieving perplexity within $4.6\%$ of the fully-synchronized baseline in $27\%$ less wall-clock time. On bandwidth-constrained links, FAUST delivers $48$--$303{\times}$ speedup over DDP.
Chat is not available.
Successful Page Load