BReD: Block Replay Dithering for Stable Low-Bit EMA Optimizer States
Heshen Zhan ⋅ Youhan Huang ⋅ Yunke Peng ⋅ Yao Wang ⋅ Linghui Kong ⋅ Ziwei Zhu ⋅ Qingyu Han ⋅ Yaoyuan Wang ⋅ Congliang Chen ⋅ Ruoyu Sun
Abstract
Low-bit storage of optimizer states can substantially reduce the memory footprint of large-scale training. In practice, however, these states are not quantized once: at every training step, they are dequantized, updated, requantized, and written back. This iterative quantization pipeline can destabilize training. We identify a failure mode behind this instability: for states updated via exponential moving averages (EMAs), quantization errors are recursively fed into subsequent updates. As a result, biased errors accumulate over time while high-variance errors are amplified, particularly when the EMA decay factor is close to 1. To address this error accumulation, we propose **BReD **(**Block **Re**play **D**ithering), a method that reduces rounding bias and controls rounding variance during quantization. BReD adapts classical subtractive dithering to optimizer-state quantization by combining deterministic seed replay with block-wise shared dither values, avoiding auxiliary random-tensor storage and per-element dither values generation. Across pretraining (five optimizers at 120M; AdamW and Muon at 1.1B and 3.4B model) and supervised fine-tuning (AdamW and Muon at 7B model), 4-bit BReD closely matches training with full-precision optimizer states, with PPL shifts $\leq0.5$ in pretraining and $\leq0.2$ in fine-tuning and average downstream performance changes within 0.7 points. Moreover, 3-bit BReD preserves stable convergence in the evaluated settings, offering a more aggressive alternative.
Chat is not available.
Successful Page Load