Fixed-Boundary Token Exploration for Parameter-Efficient Continual Pre-Training of Mixture-of-Experts
Abstract
Continual pre-training of mixture-of-experts (MoE) language models must adapt token--expert assignments to new data while preserving pretrained capabilities. Although experts can learn transferable features, the routing boundaries formed through pre-training may not transfer to shifted token representations. We first investigate whether tuning the router alone can recover suitable assignments. With the backbone and experts frozen, replacing the linear router with increasingly wide nonlinear routers yields limited gains and does not eliminate forgetting. We then propose Fixed-Boundary Token Exploration (FBTE), which freezes the pretrained router and experts and learns a zero-initialized residual transformation of token representations. Tokens can retain their original assignments or cross fixed routing boundaries to explore alternative experts. On OLMoE-1B-7B with Dolmino-Mix continual pre-training dataset, FBTE uses the same trainable-parameter budget as router tuning, achieving 8.22\% lower evaluation perplexity than router tuning and 83.78\% lower average forgetting across four evaluation domains. These results favor token exploration over tuning a shared router for parameter-efficient continual pre-training.