Why DiT Models Underperform as Representation Learners without Long Skip Connections
Abstract
Diffusion feature — intermediate activations extracted from diffusion model backbones — has emerged as a promising approach for unifying generation and vision understanding (discrimination). Despite the advancements of diffusion backbones into Diffusion Transformers (DiT), the knowledge of diffusion feature largely remains in U-Net architectures. Hence, in this work, we aim to extend the study of diffusion feature onto DiT backbones. We find that simply applying the previous methodology of diffusion feature to DiT models yields unsatisfying results. We further show that long skip connections (LSCs) are a key missing factor behind this discrepancy and propose a mechanistic explanation for how they affect representation quality. Specifically, LSCs can provide shortcuts to help noises bypass the section of the backbone where the better features reside. This hypothesis is validated using both mutual information measuring for correlation evidence and a controlled study for causal evidence. Of the two, the controlled study demonstrates that LSCs work by transporting noise information rather than just any information, indicating that LSCs enable DiT backbones to adopt a more optimized behavior pattern. Thus, our findings suggest that representation quality in diffusion models is strongly influenced by how information is routed within the backbone.