Rare-Class Signal Suppression in Long-Tailed Multi-Expert Fine-Tuning
Abstract
Multi-expert ensemble frameworks have achieved strong performance on long-tailed visual recognition by training heterogeneous expert heads on a shared backbone, but their gradient coordination strategies are inherited from balanced multi-task learning, where resolving directional conflict is the canonical challenge. We show that head-class bias acts sequentially across three pipeline interfaces, namely loss geometry, gradient coordination, and inference fusion, forming a suppression cascade in which rare-class gradient energy can be erased before it influences model parameters, a failure mode that direction-focused analysis cannot detect. A controlled study reveals that direction-modifying and simple-aggregation solvers are statistically equivalent once rare-class source signal is amplified, while energy-balancing objectives erase this amplification entirely. To break this cascade, we propose Gale (Gradient-energy Allocation for Long-tail Ensembles), a framework that applies gradient-energy allocation at every cascade stage, amplifying tail-class source signal through frequency-calibrated loss geometry, preserving it through a non-penalizing coordination rule, and routing predictions through class-specific expert competence. On CIFAR-100-LT (ρ=100), Gale achieves 59.2% overall accuracy and 46.1% few-class accuracy across four seeds (σ=0.23%), surpassing the prior best by +6.2%, with consistent improvements on ImageNet-LT and iNaturalist 2018. Code is available at Supplementary Material.