MEME: Lightweight Hierarchical Mixture-of-Experts for Unified Affective Computing
Yinan Zhang ⋅ Haoyu Zhang ⋅ Tianshu Yu
Abstract
Multimodal sentiment analysis (MSA) and emotion recognition (ER) are closely related affective understanding tasks, yet they are usually studied separately due to differences in label space, task formulation, dataset distribution, and modality dependence. In this paper, we propose MEME ($\textbf{M}$ixture-of-$\textbf{E}$xperts for $\textbf{M}$ultimodal sentiment analysis and $\textbf{E}$motion recognition), a lightweight hierarchical MoE framework for unified multimodal affective computing across MSA, conversational emotion recognition, and dynamic facial expression recognition. MEME operates on frozen text, visual, and audio features, compresses variable-length modality sequences with learned-query attention pooling, and processes them with a shared hierarchical MoE backbone. Each block first applies hard modality experts for modality-specific refinement and then task-conditioned cross experts for task-aware multimodal interaction. A $\texttt{[TASK]}$ token provides an explicit task anchor for routing, while modality dropping regularizes unified training. We jointly train MEME on nine affective benchmarks and evaluate all datasets using a single composite-best checkpoint without dataset-wise adaptation. Extensive experiments show that MEME outperforms strong task-specific baselines and recent large-model-based unified methods on most benchmarks, while maintaining favorable efficiency, robustness and generalization.
Chat is not available.
Successful Page Load