Exact Channel Decoupling via Joint Diagonalization and Uniform Splicing for Diffusion Transformer Quantization
Bingyao Yu ⋅ yijin liu ⋅ Qinkai XU ⋅ Li Li ⋅ Yuxiang Fu
Abstract
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture in high-fidelity generative modeling, but their massive computational demands and multi-step inference process limit their deployment in practical scenarios. Post-training quantization (PTQ) offers a promising solution for accelerating inference, but DiTs suffer from severe activation outlier issues, which can lead to catastrophic accuracy loss in integer formats. Although some recent outlier mitigation methods employ orthogonal transformations or diagonal scaling to smooth the activation distribution, they either solve the transformation matrix solely for activations or fail to fully balance the inter-channel variance between activations and weights. Additionally, the construction of the calibration set by taking the weighted average of the activation values for each time step introduces spurious temporal correlations. To address these issues, we propose a joint diagonalization of the activation-weight covariance matrix based on the solution of a generalized eigenvalue problem (GEVP). We also transform the feature fusion of the calibration set's temporal dimension into probabilistic sampling and assembly in the spatial dimension using Uniform Token Splicing (UTS). Specifically, compared to suboptimal baseline methods, our approach reduces the FID by 16.06 and 6.49 on both the W3A4-quantized DiT-XL/2 and PixArt-$\Sigma$ models, respectively. This demonstrates that our method successfully suppresses outliers in activations and weights, maintaining DiT's performance under low-bit quantization conditions. Code is available at the anonymous repository \url{https://anonymous.4open.science/r/JDUS-DiT-74E1}.
Chat is not available.
Successful Page Load