Advanced Routing as Regularization Allocation for Efficient Diffusion Transformer Training
Qin MA ⋅ XIAOQI SUN ⋅ bo li ⋅ Yuquan Zhou ⋅ Weizhong Zhang
Abstract
Diffusion Transformers (DiTs) have achieved strong performance in visual generation, but their dense token processing leads to slow training convergence and high computational cost. Recent token routing methods such as TREAD accelerate training by allowing a random subset of tokens to bypass intermediate layers. It can be expected that carefully tuning the routing ratios over time steps or layers can achieve more significant accelerations. However, the theoretical mechanisms underlying stochastic routing remain underexplored, making the tuning of routing ratios intractable. In this paper, we first theoretically show that stochastic token routing can be interpreted as an implicit route-sensitivity regularizer. Under this view, the routing ratio determines the strength of the induced regularization: larger routing ratios introduce stronger route-induced perturbations and therefore stronger regularization pressure. This further motivates a noise-conditioned routing strategy. In diffusion training, samples at larger timesteps contain heavier noise and often provide less stable gradient signals, suggesting that they may require stronger regularization. We therefore assign larger routing ratios to higher-noise samples, allowing the routing-induced regularization strength to adapt to the noise condition. Based on this insight, we propose \emph{Noise-Aware Routing}, which partitions samples within each mini-batch into low-, middle-, and high-noise regions and assigns progressively larger routing ratios to them. This simple noise-conditioned allocation synergistically bridges computational savings with enhanced model priors. Extensive experiments on class-conditional ImageNet generation show that our method accelerates training convergence and improves generation quality. Noise-Aware Routing achieves up to \textbf{15.6}$\times$ and \textbf{20}$\times$ convergence speedups over DiT baselines—and outperforms TREAD—on ImageNet-256 and ImageNet-512, respectively, measured by the training iterations required to match our 400K-step FID. In the guided setting, our method achieves an FID of \(2.67\) with a substantially shorter training iterations without architectural changes.
Chat is not available.
Successful Page Load