SAME: Stability-Aware Embedding Extraction in Mixture-of-Experts Language Models
Abstract
Extracting sentence embeddings from MoE-based large language models is a promising direction, as they provide greater model capacity than dense models at comparable computational cost. Existing works perform the standard forward propagation process to extract embeddings, overlooking a key stability requirement in the MoE encoding process: semantically similar texts should be encoded into similar representations. In this work, we identify two key stability-related phenomena in MoE models: (1) routing stability varies across layers, and (2) shared and routed experts exhibit different levels of stability. To address these issues, we propose SAME, a Stability-Aware MoE Embedding extraction framework that dynamically allocates activated experts across layers and adjusts expert outputs to mitigate component instability and improve embedding quality. Specifically, SAME assigns fewer activated experts to layers with less stable routing, where layer-wise routing stability is measured by the average overlap between the experts activated before and after injecting Gaussian noise into each layer’s router inputs. Additionally, SAME reduces the contributions of routed experts to alleviate the impact of their instability. Notably, our method is training-free, seamlessly integrates with existing approaches, and incurs no additional inference overhead. Experiments on semantic textual similarity benchmarks demonstrate that SAME consistently improves the quality of extracted embeddings across multiple MoE backbones.