Structure, Subspace and System: Push the Real Limit of Extremely Low-Bit Quantization for MoE-LLMs
Jiaqi Zhao ⋅ Zichen Li ⋅ Miao Zhang ⋅ Yixuan Dong ⋅ Weili Guan ⋅ Liqiang Nie
Abstract
Mixture-of-Experts (MoE) large language models (LLMs) suffer significant inference and storage overheads, which makes extremely low-bit quantization (sub 2-bit) highly desirable. However, existing methods for dense LLMs typically quantize each expert independently and require it to approximate the original mapping over the entire input space. We argue that this formulation is overly conservative and structurally mismatched for MoE: 1) experts in the same MoE layer are not fully independent but operate on a shared hidden representation space. Under low-bit budgets, independently binarizing each expert will distort the activation-sensitive input geometry, which leads to severe performance degradation, and 2) due to routing-induced specialization, each expert only serves a concentrated subset of tokens, whose activations typically occupy a much narrower subspace than that of dense layers. More importantly, when the weights approach 1-bit, the dominant deployment bottleneck is no longer the weights themselves, but the metadata (e.g., scaling factors, bitmap, grouping index) which can raise the real effective inference precision to around 3-bit. This issue, however, has largely been overlooked in prior works. Building on these perspectives, we introduce a novel extremely low-bit quantization framework for MoE-LLMs, called TriS-MoE, which considers Structure, Subspace and System. Specifically, we first extract a shared high-precision input-side backbone to preserve activation-sensitive input geometry, while binarizing the remaining components to save costs. Secondly, we propose to redirect expert quantization errors into the null space of routed activations to minimize output perturbation on the subspace actually used by each expert. Finally, we co-design a bitmap compression method based on Golomb-Rice coding which approaches the Shannon entropy lower bound and a specialized streaming loading mechanism to reduce metadata overhead during inference. Extensive experiments on MoE-LLMs demonstrate our TriS-MoE outperforms the strongest baseline by 8.31% average accuracies, and also achieves an average 6.32$\times$ reduction in inference memory on real systems, enabling Qwen3.5-35B-A3B and Mixtral-8$\times$7B to run on a single consumer-GPU.
Chat is not available.
Successful Page Load