Comparing Synchronous and Asynchronous Pipeline Parallelism Across Training Configurations
Junkyu Park ⋅ Dongyeop Lee ⋅ Namhoon Lee
Abstract
Pipeline-parallel training can coordinate optimizer updates synchronously by draining in-flight microbatches or asynchronously by updating stages while the pipeline remains active. Although asynchrony can improve token throughput, asynchronous training may require more tokens to reach a given quality, leaving unclear which schedule minimizes time to target. We compare two one-forward/one-backward (1F1B) schedules---Synchronous 1F1B-Flush and Asynchronous 1F1B-Async---for AdamW, Muon, and SOAP at pipeline degrees 8 and 32 on a GPT-style language model, with matched training conditions within each optimizer. We combine measured loss-versus-token trajectories with a profile-driven reconstruction of one-GPU-per-stage runtime, validated at P8 and used to project P32 timing. At P8, Asynchronous execution achieves a 2.6--5.2$\times$ speedup in fixed-token-budget runtime, while Synchronous execution reaches the same optimizer-specific loss target with fewer tokens for all three optimizers. Consequently, Synchronous execution is 1.7$\times$ faster to target for AdamW, whereas Asynchronous execution is 3.3$\times$ and 2.4$\times$ faster for Muon and SOAP, respectively. Under the P32 projection, Asynchronous fixed-token-budget runtime decreases relative to P8, yet its time to target increases for all three optimizers. Thus, neither synchronization convention nor throughput alone determines the faster schedule; schedule selection should be based on end-to-end time to the relevant quality target under the complete training configuration.
Chat is not available.
Successful Page Load