S²MoE: Shared-Subspace Mixture of Sparse Experts
Abstract
Model merging offers a training-free paradigm for integrating multiple specialized models into a unified multi-task system, enabling efficient knowledge reuse from models. However, parameter interference remains a critical challenge, limiting the scalability and efficacy of merged models. SVD-based merging techniques represent the current state-of-the-art, excelling in alleviating interference through task-specific knowledge extraction and subspace orthogonalization. % due to aggressive parameter truncation Despite these advances, they inevitably incur performance degradation due to aggressive parameter truncation, which discards valuable information and hinders inference accuracy. Our analysis reveals this truncation-induced knowledge loss as the fundamental bottleneck, motivating the need for mechanisms that preserve and recover discarded expertise without retraining. To address this, we introduce S MoE, a novel framework that augments SVD-based merging with a mixture of sparse experts to refine the truncation gap. S MoE comprises three key components: (1) a shared module that aggregates multi-task knowledge, (2) sparse experts that selectively recover truncation knowledge, and (3) a subspace mapping router to achieve precise training-free routing. Extensive experiments across vision and language benchmarks demonstrate up to 1.6% and 1.2% performance gains over SOTA methods, effectively validating robust and scalable model merging.