Comp$^2$VLM: A Hybrid Framework Combining Quantization and Lossless Compression for Efficient Vision-Language Models
Abstract
Vision–language models (VLMs) exhibit strong multimodal reasoning, yet their growing parameter counts incur prohibitive memory footprints and inference latency, hindering real-time deployment. Post-training quantization alleviates these costs, but aggressive low-bit quantization often degrades accuracy by disrupting visual–linguistic alignment. Meanwhile, lossless compression preserves accuracy but suffers from metadata overhead and sequential decoding, limiting GPU parallelism. We propose Comp²VLM, a hybrid framework that jointly designs quantization and lossless compression for efficient VLM serving. We introduce Duplication-Enhanced Quantization (DEQ), which adaptively selects group-wise scales to reduce the entropy of quantized representations while preserving reconstruction fidelity, thereby improving compressibility. To realize runtime gains, we present Quantized Data-aware Lossless Compression (QPress), a hardware-friendly scheme based on a universal Huffman table that enables block-wise metadata sharing and GPU-parallel decoding. Across diverse VLMs and benchmarks, Comp²VLM maintains competitive accuracy while reducing the effective precision to W3.5A3.2~4.9. On LLaVA-onevision-7B, it achieves +4.7% TextVQA and +7.7% OCRBench accuracy over W4A8-based SOTA at an average bit-width of W3.5A4.6. DEQ yields up to 33% entropy reduction, and QPress improves compression efficiency by up to 13.3% with only ~4% end-to-end latency overhead. To our knowledge, Comp²VLM is the first unified framework to co-design quantization and lossless compression for VLMs, demonstrating their complementary benefits beyond standalone methods.