LEGO: Sizing Rules for Budget-Aware Dense-to-MoE Conversion of Vision-Language Models
Qishen Yin ⋅ ZiangWu ⋅ Juntong Wu ⋅ Peng Jin ⋅ Li Hao ⋅ Tanghui Jia ⋅ Liuhan Chen ⋅ Bin Zhu ⋅ Li Yuan
Abstract
We propose LEGO, a budget-aware Dense-to-MoE methodology that converts pretrained vision-language models into efficient FFN-MoE variants under a target architectural budget $((P_{total},P_{active}))$. Our study is motivated by a counter-intuitive recovery pattern: fixed activation ratios that work well for some dense backbones can degrade sharply on others under the same recovery recipe, indicating that raw sparsity alone is an unreliable design rule. Through controlled recovery experiments across VLM families and scales, we identify the activated-FFN-to-hidden ratio $(D_I/D_H)$ as a backbone-comparable structural factor and summarize empirical sizing rules for avoiding severe bottlenecks while limiting diminishing returns. LEGO operationalizes these rules through a discrete budget-aware search over Split, Upcycle, and Hybrid constructions, and introduces a moment-matching scaling factor to reduce the initialization scale shift caused by sparse FFN aggregation. Finally, we use a two-stage multimodal recovery recipe that first stabilizes the sparse LLM with the vision tower frozen, then performs joint fine-tuning to recover the full VLM. Under matched data and training protocols, LEGO improves Dense-to-MoE recovery quality and turns costly blind architecture search into a small set of rule-guided candidates. Codes can be found in supplementary materials.
Chat is not available.
Successful Page Load