GLUQuant: Leveraging Expert Activation Geometry for 2-Bit Diffusion MoEs
Honjar Xing
Abstract
Diffusion language models denoise a whole canvas of positions at once, and the largest open ones are mixture-of-experts, which puts almost all of their weight in the expert tensors. In DiffusionGemma-26B-A4B they hold 90.4% of the text parameters and 45.7 of its 50.5 GB, so the experts are the memory bottleneck and quantising them is the compression that matters. That compression pays off up to a certain point. A roofline over the positions a denoising step presents bounds the width worth quantising to, since below it smaller tensors buy no further speed. For this model the bound is 2.07 bits, 26× higher than for an autoregressive decode, so what a diffusion MoE needs is a good two-bit quantiser. Our main finding is that those two bits should not be spent evenly, because the two matrices of a GLU expert read opposite activation geometry. gate_up reads a subspace shared by every expert in its layer, while down_proj reads one private to each expert, overlapping no more than unrelated random subspaces do. The GLU causes this and not the router, and the same split appears in an autoregressive MoE. Shared structure amortises over a layer and private structure cannot, so a private direction costs three times a shared one. We propose GLUQuant, which gives each matrix the basis its own geometry calls for and codes the residual with an $E_8$ nested lattice that carries no side information at all. At 2.02 bits and without fine-tuning it compresses the experts 7.91×, brings the model to 10.6 GB, and runs the expert path 1.76× faster than bf16. It beats 2-bit GPTQ and expert-level mixed precision on all 8 tasks we score and a distillation-free QuIP# on 6, at fewer bits than any of them. On GSM8K it stays within 0.9% of dense. The format stores no calibrated state, so swapping the calibration corpus moves it 0.15 points where GPTQ moves 11.15.
Chat is not available.
Successful Page Load