SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
Boyao Wang ⋅ Zhihan Lei
Abstract
Modular networks pursue specialization through learned routers, gates, and load-balancing losses. However, at matched total-parameter budgets, learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training-signal granularity and the target categories? We show that when each training and inference unit carries one coarse category tag, a fixed parameter-free routing scheme (SpecDrop) induces branch-category specialization and matches or exceeds the learned-routing baselines we evaluate at matched parameter budget; when this alignment breaks, the scheme matches the multi-branch No-Routing baseline. SpecDrop assigns each of $K$ branches weight $p_{\mathrm{a}}$ for its category and a small leakage $p_{\mathrm{i}}{>}0$ otherwise, merged through a category-independent fixed denominator that calibrates merged-branch magnitude to single-branch scale, with no learnable routing parameters or auxiliary losses. Category labels are required at inference, drawn from dataset metadata. On vision tasks where each image has one superclass label, SpecDrop reaches $\mathbf{79.23\%}$ on CIFAR-100 with ResNet-110, outperforming dense by $+4.75$, and $\mathbf{79.89\%}$ on ImageNet-1K with Vision Transformer ViT-S/16, outperforming the matched-supervision No-Routing baseline with a shared expert by $+6.53$. SpecDrop reaches the highest top-1 among the multi-branch routing baselines we evaluate at this parameter budget (Soft MoE and Mod-Squad on ImageNet). On NLP tasks where training units span multiple categories, SpecDrop matches the multi-branch No-Routing baseline on both SlimPajama-6B language modeling with a 30M Transformer and SuperNI instruction tuning over Llama-3.2-1B with LoRA. Furthermore, SpecDrop matches or outperforms multi-branch routing baselines including Demix and LoRAMoE, and is statistically tied with HydraLoRA. Per-branch pruning sensitivity reveals four specialization regimes ranked by category clarity: strongest on ImageNet-1K, strong on CIFAR-100, weakening on SlimPajama, and anti-aligned on SuperNI LoRA. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://anonymous.4open.science/r/C30862.
Chat is not available.
Successful Page Load