What MoE Routing Controls: Causal Interventions on Expert Trajectories
Abstract
In mixture-of-experts (MoE) language models, the sequence of experts a token visits across layers correlates with its semantic function, but whether this routing causally determines behavior is untested. We define a target-expert trajectory by taking the modal top-1 expert over concept or role exemplars, and intervene by adding a sustained bias to that one expert's router logit at each layer during generation. In Qwen3-30B-A3B (top-8 routing) this promotes the target expert into each token's selected set. We show that this intervention alone biases the model's output toward the concept associated with that path, causal evidence that routing shapes the output distribution rather than merely correlating with it. We then test how far this control reaches and how much to trust it, with controls that rule out generic corruption, injected vocabulary, and gate-weight inflation, and paired significance tests and confidence intervals throughout. The reach is limited, as the intervention changes generated text for some targets but only shifts the next-token distribution for others. Format interventions can produce large effects in the generated text, whereas concept interventions are selective. We also show that a trajectory built from Chinese religion tokens raises English religion-word probability, which rules out surface copying, but the absolute probabilities stay small and the generated text barely changes. Which concepts move at all is a robust ordering, and we test two explanations for it, embedding distinctiveness and lexical frequency, and find neither supported.