SACQ: Structure-Aware Codebook Quantization for Sub-2-Bit Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models produce remarkably capable robot behavior, but their memory footprint makes deployment on resource-constrained edge platforms impractical. Weight quantization is the natural remedy, but post-training quantization (PTQ) recipes inherited from large language models degrade sharply below two bits. These recipes rest on two assumptions that do not hold for VLAs. First, they apply one objective to every projection, ignoring that each role feeds a different downstream computation. Second, they allocate precision without identifying which layers actually matter for the executed action. We introduce SACQ, a sub-2-bit weight-only vector-codebook framework for VLAs, addressing these in turn: (1) structure-aware codebook fitting, which fits each layer’s codebook under a distortion geometry derived from that projection’s role rather than Euclidean weight distance; and (2) action-aware mixed-bit allocation, which ranks layers by how much the decoded action degrades when each is given fewer bits. A final calibration stage then refines codebooks and scales against the decoded action, with assignments frozen. On GR00T N1.7 across all four LIBERO suites, SACQ retains BF16-level closed-loop success at 1.40 bits per candidate weight under matched three-seed evaluation, i.e., 94.42±0.80% versus 94.67±0.45%, a mean paired difference of −0.25 points, while reducing the deployed model from 5.677 to 0.969GB.