Mixture-of-Hierarchical Experts: Optimized Mamba Architecture for Vision Diffusion
Yejun Jung ⋅ Dongyun Kim ⋅ Jinsun Park
Abstract
Existing Mamba-based diffusion backbones struggle to achieve competitive performance in vision tasks, as purely sequential propagation lacks an explicit spatial inductive bias and relies on a fixed computation pattern across timesteps. We address this limitation by proposing \textbf{Mixture-of-Hierarchical Experts (MoH)}, a Mamba-based diffusion backbone that performs timestep-conditioned routing over architectural components. MoH dynamically selects representation levels by its hierarchical expert space, adapts feature transformations, and adjusts Mamba scan depth according to the denoising state, enabling flexible modeling of spatial structure and multi-scale dependencies. This design leads to substantial performance gains. On unconditional CelebA-HQ $256 \times 256$, MoH reduces FID from 14.27 to 7.99, and on MS-COCO, it improves FID from 41.80 to 20.00, significantly outperforming prior Mamba-based models while improving training efficiency. These results suggest that routing over architectural components can provide an effective mechanism for improving Mamba-based diffusion backbones.
Chat is not available.
Successful Page Load