ProQuant: Progressive Quantization-aware Training for Edge MLLMs
Yufei Xue ⋅ Yushi Huang ⋅ Jiawei Shao ⋅ Pingcheng Dong ⋅ Yonghao Tan ⋅ Shiyao Li ⋅ Kwang-Ting Cheng ⋅ Jun Zhang
Abstract
Multi-modal large language models (MLLMs) with impressive perception and understanding capabilities will enable a range of downstream applications. However, their substantial computational and memory requirements hinder the real-world implementation, particularly for resources-constrained edge devices. Quantization has emerged as an effective solution for deploying large models. While post-training quantization (PTQ) successfully retrains the performance at moderate bit-width (\textit{e.g.}, W8A8), it suffers from severe performance degradation under low-bit quantization (\textit{e.g.}, W4A8/W8A8), especially for small-scale models that are more suitable for edge deployment. In this paper, we propose \texttt{ProQuant}, a novel weight-activation quantization-aware training (QAT) framework tailored for low-bit MLLMs. First, to align the limited finetuning dataset distribution with the pretrained dataset distribution, we propose a self-relabeling (SRL) scheme to resample the responses of the open-source multimodal dataset. Second, we are the \textit{first} to take a decoupled view of multi-modal (MM) understanding and language modeling abilities. We propose a progressive block activation (PBA) mechanism to prioritize the MM understanding recovery. We also introduce a vision-sensitive factor $\vec{\tau}$ that effectively identifies the MM-critical layers. Extensive experiments across various benchmarks, covering models with 2B$\sim$8B parameters, show that \texttt{ProQuant} significantly outperforms existing methods. For example, our W4A8 \texttt{Qwen3-VL-2B-Instruct} improves average accuracy by impressively 6.6\% and achieves performance comparable to full precision counterparts with only $\sim$1\% performance drop. Code will be released upon acceptance.
Chat is not available.
Successful Page Load