Fewer Expert Transfers Without Changing Routing: Expert Eviction for MoE Offloading
Pierre-Luc Bacon ⋅ Diego Calanzone ⋅ Jean-Maxime Larouche
Abstract
Large mixture-of-experts (MoE) language models often keep only part of their expert weights on the GPU and copy the rest from CPU memory when needed. Recent learned methods try to reduce this cost by keeping the same experts active for several tokens, but fewer routing changes need not produce fewer copies, and forcing a fixed set can change the model's predictions. We compare these methods with the standard least-recently-used (LRU) cache in a harness that performs and times real CPU-to-GPU transfers. On an L40S, every tested method that runs faster than LRU increases next-token loss by at least $0.067$ nats; a model based on measured copy time also selects the fastest policy when transfers are made $10\times$ and $50\times$ slower. Changing only which expert the cache evicts, while leaving routing untouched, cuts transfers by $2.1\times$ with no quality loss and comes within $3.5\%$ of the best possible schedule. Sixteen tokens of future demand recover $96\%$ of this gain, while a frequency-based rule that uses only past requests recovers $27\%$.
Chat is not available.
Successful Page Load