ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
Arash Akbari ⋅ Arman Akbari ⋅ Masih Eskandar ⋅ Qitao Tan ⋅ Yixiao Chen ⋅ Jingwu Luo ⋅ Bertha Pangaribuan ⋅ Liyun Zhang ⋅ Jennifer Dy ⋅ Geng Yuan ⋅ Xue Lin ⋅ Gaowen Liu ⋅ Stratis Ioannidis ⋅ Yanzhi Wang
Abstract
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute and memory demands strain deployment on edge robotic platforms. Aggressive weight quantization (e.g., sub-4-bit) is the natural lever, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce \textbf{ActQuant}, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm. On the LIBERO benchmark, ActQuant achieves state-of-the-art success rates across the sub-4-bit regime, retaining 94.95\% on OpenVLA-OFT and 94.8\% on $\pi_{0.5}$ at 3.0 bits-per-weight (vs. 96.9\% and 97.0\% for the half-precision baselines, respectively), while reducing backbone memory by up to $5.3\times$; on the physical robot, ActQuant successfully maintains the baseline's success rate. To our knowledge, ActQuant is the first PTQ approach that can quantize VLA models below four bits while retaining practical task success with very limited calibration data.
Chat is not available.
Successful Page Load