vExpert: Virtualizing Expert Storage for Adaptive Load Balancing in Distributed MoE Inference
Wenxun Wang ⋅ Xiuhong Li ⋅ Yida Wang ⋅ Chen Tang ⋅ Zongle Huang ⋅ Ke Hong ⋅ Yu Wang ⋅ Yongpan Liu
Abstract
Mixture-of-Experts (MoE) models have become prevailing, yet their distributed deployment suffers from severe dynamic load imbalance due to the conflict between static expert placement and inherent routing dynamism. While host DRAM offers capacity to avoid static device storage for load balancing, we identify existing offloading works fundamentally fail in distributed inference due to two overlooked challenges: (1) coupled performance impacts of runtime Host-to-Device transfers due to load dynamism and shared PCIe topology; (2) system-level memory overhead incurred by DRAM usage at scale. In this paper, we propose vExpert that shifts DRAM offloading paradigm from partial residency to full redundancy by virtualized expert storage. We pioneer the first systematic performance model that jointly captures H2D latency, PCIe contention, and load imbalance, building a lightweight allocator to dynamically map physical experts to virtual slots. A novel disaggregated manager is designed to reduce memory overhead in multi-process environments. Evaluated on DeepSeek-V3.1 within an 8-node H100 cluster, vExpert improves load balance ratio by up to 55\% and achieves 1.2$\times$ end-to-end speedup. Code is available at https://anonymous.4open.science/r/vExpert-B743/.
Chat is not available.
Successful Page Load