ACQueReLlo: Alignment-Aware Constrained Quantization via Reinforcement Learning for Large Language Models
Abstract
Post-training quantization is a standard tool for the efficient deployment of large language models (LLMs). In safety-aligned models, it introduces a critical and underexplored risk: preserving utility under compression does not ensure that safety alignment is retained. Small perturbations from quantization can disproportionately degrade refusal to harmful prompts, leading to increased compliance with harmful prompts or over-refusal of benign prompts, even when aggregate performance appears stable. This gap reveals a fundamental limitation of existing quantization approaches, which primarily optimize for utility while overlooking failures due to safety alignment degradation. This paper presents ACQueReLlo, a reinforcement learning framework for mixed-precision post-training quantization. It formulates block-wise bit allocation as a constrained Markov decision process and learns a quantization policy under explicit constraints on utility, harmful prompt refusal, benign prompt over-refusal, and overall bit budget. To ensure tractable training, the policy relies on proxy evaluations over sampled subsets of utility, harmful, and benign prompts. Experiments on various LLMs show that the learned policy achieves a better overall balance between compression, utility, and safety alignment than uniform quantization and other baselines.