Error Feedback Enables Activation Compression for Pipeline-Parallel Diffusion Language Model Inference
Jinseok Chung ⋅ Namhoon Lee
Abstract
Pipeline-parallel diffusion language model inference sends large activations across stage boundaries at every denoising evaluation, making communication a major bottleneck. Aggressive compression reduces this traffic, but its errors propagate through later stages and denoising evaluations, degrading task performance. Error feedback is a natural way to preserve information lost during compression. However, applying it to activations is not straightforward: unlike a gradient update, an activation is consumed immediately and has no natural future update that can carry its error. We show that the denoising loop provides this opportunity: repeated evaluations produce related activations at each pipeline boundary, so information omitted from one transfer can be included in later ones. This enables error feedback without extra communication while model weights remain partitioned across pipeline stages. Across quantization, top-$k$ sparsification, and low-rank compression, our method cuts communication by up to $156\times$ while matching uncompressed task performance; without error feedback, benchmark scores can fall close to zero. We also demonstrate substantial latency reductions in a real eight-GPU pipeline.
Chat is not available.
Successful Page Load