FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
Abstract
Parameter-efficient fine-tuning (PEFT) has emerged as a crucial paradigm for adapting large language models (LLMs). However, standard PEFT methods often struggle in multi-task fine-tuning settings due to task interference and a limited parameter budget. Recent approaches incorporate the mixture-of-experts (MoE) architecture, referred to as mixture-of-parameter-efficient-experts (MoPE), to alleviate this issue by dynamically routing inputs to specialized experts. However, these methods remain based on spatial parameterization, which may introduce structural redundancy and additional parameter overhead. To address these limitations, we revisit model adaptation from a spectral perspective. Our analysis uncovers heterogeneity in frequency sensitivity across model layers and downstream tasks, indicating that adaptation should be frequency-aware rather than uniformly parameterized in the spatial domain. Motivated by this insight, we propose FourierMoE, a novel framework that unifies spectral parameterization with MoE via frequency-specialized experts and the learning of conjugate-symmetric complex coefficients. We conduct extensive experiments across multiple model families on a diverse range of tasks, including commonsense reasoning, math reasoning, image classification, and natural language understanding. Experimental results on 28 benchmarks show that FourierMoE consistently outperforms full fine-tuning (FFT) and 15 competitive baselines in both single-task and multi-task settings, while requiring significantly fewer trainable parameters, demonstrating superior effectiveness, efficiency, and adaptability.