FineRMoE: Dimension Expansion for Finer-Grained Expert with an Upcycling Approach
Abstract
As revealed by the scaling law of Mixture-of-Experts models, model performance ceases improving once the granularity of the intermediate dimension exceeds the optimal threshold, limiting further gains from single-dimension fine-grained design. To address this bottleneck, we propose FineRMoE (FineR-Grained MoE), an architecture that extends fine-grained expert design to both intermediate and output dimensions, aiming to enhance expert specialization beyond the single-dimension limit. We further introduce a bi-level sparse forward computation paradigm and a specialized routing mechanism to govern the activation. In addition, to obviate the prohibitive cost of pre-training from scratch, we devise a generalized upcycling method to build FineRMoE in a cost-effective manner. Extensive experiments performed under the upcycling paradigm demonstrate the superior performance achieved by FineRMoE across ten standard benchmarks. Compared with the strongest upcycling baseline method, FineRMoE achieves significant parameter and inference efficiency gains.