AdERA: Adaptive Exponent Reuse for Lossless Allgather in Sharded MoE Training
Abstract
In training Mixture-of-Experts models, sharded data parallelism splits each expert’s parameters across GPUs. Before each layer runs, GPUs must use Allgather to rebuild the full weight matrix. This communication can take a large part of each training iteration. Prior work reduces this cost with lossy compression, but lossy methods can affect accuracy. We propose Adaptive Exponent Reuse Allgather, or AdERA, a lossless Allgather compression method for sharded MoE training. Our key observation is that, after a short warmup phase, most weight exponents stay the same across iterations. AdERA stores these exponents locally and sends only the sign and mantissa when the exponent has not changed. The receiver then rebuilds the exact original weight by combining the cached exponent with the received sign and mantissa. Because parameter matrices have different shapes across layers, AdERA compresses only when compression saves time. It also overlaps compression with Allgather communication and computation. Our real experiments and large-scale trace-driven simulator show that, compared with the lossless baseline, AdERA achieves a 3.70× speedup on 16 GPUs and is projected to reach 4.28× on 128 GPUs, while preserving bitwise-exact parameter reconstruction. Compared with the lossy baseline, AdERA achieves a 3.93× speedup on 8 GPUs and is projected to reach 4.42× on 128 GPUs. The source code is publicly available.