What is the role of Learning Rate in a Mixture-of-Experts ResNet
Abstract
A large learning rate is one of the few reliably beneficial choices in deep learning, and in a Mixture-of-Experts network it has another dimension to act on: not only the weights, but the discrete decision of which expert each input is sent to. We ask what it does to that decision, and whether the two effects are connected. Varying the learning rate in a three-site top-1 MoE ResNet on CIFAR-10, across three optimisers and the full range of load-balancing strengths, we find that a larger step helps along every axis we can measure --- clean accuracy, corruption robustness, and the stability of the routing itself. From early stage of training, we notice a collapsed router independent of the step size, yet by using a load balancing techniques, appropriately large learning rate is the gain that decides whether that loss can act. Pushing the balancing strength well past that point does balance the routing equally, but at the cost of router fragility. We show that it happens both along the trajectory and in the presence of corruptions; with small learning rate loses most, which connects the routing account to the generalisation performance. All reported results reached zero training error and validated with multiple seeds.