LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
Yann Bouquet ⋅ Alireza Khodamoradi ⋅ Sophie Y Shen ⋅ Kristof Denolf ⋅ Mathieu Salzmann
Abstract
Post-training quantization (PTQ) of large diffusion transformers degrades generative quality at 4-bit quantization. Low-rank approximation methods are a promising solution and append auxiliary linear branches to restore performance. Current state-of-the-art approaches keep these branches at high precision (W16A16) and rely on heavy, data-dependent calibration for initialization. Their branch is defined as a precise low-rank approximation of a full-rank matrix and therefore cannot keep precision at sub-16 bit levels. We challenge both limitations with LoRaQ (Low-Rank Approximated Quantization), a data-free calibration approach that optimizes quantization error compensation. By overcoming the need for high-precision branches, LoRaQ enables the first fully sub-16 bit pipeline, allowing the low-rank branch itself to be quantized. We demonstrate that, at equal memory overhead, LoRaQ outperforms the state-of-the-art methods in their native implementations on Pixart-$\Sigma$, SANA and Flux.1. We also analyze mixed-precision configurations, showing that setups such as W8A8, W6A6, and W4A8 for the low-rank branch, alongside a W4 residual layer, yield superior results while maintaining a fully quantized architecture compatible with modern mixed-precision hardware, and translate into kernel-level speedups of up to $3.19\times$ on AMD MI355.
Chat is not available.
Successful Page Load