Cache-Aware Routing: Towards Efficient and Stable MoEs with Reinforcement Learning on Hardware Features
Abstract
Mixture of Experts (MoE) can be made more efficient by changing the routing patterns: increasing expert re-use reduces the number of sub-networks to activate, allowing for tweaks to increase hardware efficiency, such as efficient expert offloading. Various proposed methods introduce a regularization objective that accounts for hardware features, such as cache simulation or modularization via hierarchical learning (options). We propose to adopt the former from a reward shaping perspective: with reinforcement learning, we can fine-tune the MoE router and experts to increase cache hits over a subset of experts. We observe that this encourages the model to change its routing patterns without shifting the semantics, as measured by minimum accuracy drop in downstream benchmarks and perplexity. Moreover, we show that cache-aware routing (CAR) can improve routing stability and expert modularity. Similarly, an option-critic-style controller achieves near-perfect routing consistency, but it does so by collapsing onto a handful of experts, coinciding with the worst robustness in our comparison. Routing stability, we conclude, is necessary but not sufficient for useful modularity: it must be checked against routing diversity and downstream robustness directly. We fine-tune 1B and 7B models and compare our method, CAR-MoE, to previously proposed objectives through semantic (math, coding, chat) and quantitative (cache hit ratio, routing stability, specialization index) tasks.