SPA-Q: Structure-Preserving Adaptive Post-Training Quantization for Monocular Depth Estimation
Jaemin Choi ⋅ Jincheol Yang ⋅ Nahyun Lim ⋅ Yun-Seong Jeong ⋅ Matti Zinke ⋅ Hyunwoo Yu ⋅ Suk-Ju Kang
Abstract
Monocular depth estimation (MDE) has advanced rapidly with the emergence of foundation models such as Depth Anything. Their transformer-based architectures provide strong generalization across diverse scenes and domains, but also incur high computational and memory cost, making efficient deployment challenging. Post-Training Quantization (PTQ) provides an efficient and practical solution for model compression, yet low-bit PTQ remains challenging for MDE models. When applying PTQ to Depth Anything, we identify two key challenges: (i) independently minimizing the quantization error of the query and key projections fails to preserve the attention maps induced by their interaction, and (ii) quantization errors accumulate across layers, resulting in intermediate feature distribution shifts. To address these issues, we propose SPA-Q, a Structure-Preserving Adaptive PTQ framework for MDE, consisting of two main components. First, Attention-Preserving Calibration (APC) determines query and key quantization parameters by matching the full-precision attention distribution. Second, Channel-Wise Distribution Alignment (CWDA) learns channel-wise affine transformations to mitigate quantization-induced distribution shifts, and the learned parameters are absorbed into the weights after training. Experimental results show that SPA-Q consistently outperforms existing PTQ methods under 4-bit quantization, achieving an average 25.9\% reduction in AbsRel and a 17.3\% improvement in $\delta_1$ across NYUv2 and KITTI datasets.
Chat is not available.
Successful Page Load