MoEDDI: Semantic Mixtures of Experts for Long-Tailed Drug–Drug Interaction Prediction
Abstract
Descriptor-based drug-drug interaction (DDI) models typically apply a single global predictor despite heterogeneous physicochemical evidence and severely long-tailed labels. We introduce MoEDDI (Mixture-of-Experts for Drug-Drug Interaction prediction), a chemically organized mixture-of-experts that routes overlapping descriptor families using low-dimensional statistics and combines their sparse top-2 mixture with a global T-DDI pathway through a residual logit correction. The residual parameterization can exactly preserve a supplied T-DDI checkpoint at initialization, although the evaluated models are trained from scratch. Across complete test partitions with 178, 86, and 92 classes, MoEDDI improves macro-F1 over reproduced seed-2026 T-DDI baselines from 0.8284 to 0.8485, 0.4729 to 0.6536, and 0.1807 to 0.4671. On the 3,780-feature corpus, the lowest-support group (89 classes with fewer than 50 test examples) shows the largest mean classwise gain (+3.30 F1 points). Calibration improves on one corpus but worsens on two. These single-seed results support semantic conditional specialization while motivating matched ablations and distribution-shift evaluation.