Diffusion Transformers with Residual Adaptive Layer Normalization
Ge Wu ⋅ Minxing Luo ⋅ Yikai Ge ⋅ Lei Wang ⋅ DanDan Zheng ⋅ Rui Liu ⋅ libin wang ⋅ Yu-Liang Zhan ⋅ Taihang Hu ⋅ Jingdong Chen ⋅ Xiang Li
Abstract
Residual connections with PreNorm are the default design in modern transformers due to their strong optimization stability, but recent research has shown that they weaken the functional differentiation of deeper blocks and increase representational redundancy in later layers. This issue also appears in diffusion models using similar architectures, where the deep block is expected to support progressive denoising refinement rather than repeated computation over similar states. Motivated by this, we revisit cross-layer shortcuts in diffusion transformers and argue that their role should go beyond merely transferring shallow information, instead guiding deeper blocks to improve depth utilization. To this end, we propose \textit{\textbf{Res}idual Adaptive Layer \textbf{Norm}alization} (\textbf{ResNorm}), which routes residual cross-layer shortcuts into the existing adaLN modulation. ResNorm preserves the optimization advantages of PreNorm residual connections, while reusing the native adaLN interface to transmit early-layer representations as modulation signals. This design naturally combines cross-layer history with the model's native condition, such as timestep and class embeddings. ResNorm allows the earlier structural representation to adaptively scale, shift, and gate deeper blocks. This not only produces more differentiated deep representations and reduces cross-layer similarity in later blocks, but also accelerates training convergence and improves performance. On ImageNet 256$\times$256, SiT-XL/2 + ResNorm achieves up to $\textbf{2}\times$ and $\textbf{36}\times$ faster training than SiT-XL/2 + REG and SiT-XL/2 + REPA. Moreover, SiT-L/2 + ResNorm trained for only 400K iterations already outperforms SiT-XL/2 + REG.
Chat is not available.
Successful Page Load