AdaqL: Adaptive Mixed-Precision Layer-wise Quantization for Large Language Models
Abstract
Mixed-precision quantization can reduce the memory cost of large language models (LLMs) by assigning different bit-widths to layers with different quantization sensitivities. However, determining an effective bit-width configuration is challenging due to the exponentially large search space, while existing sensitivity-based methods rely on local proxies that may not fully capture the behavior of the jointly quantized model. We propose AdaqL, which instead formulates mixed-precision allocation as an end-to-end optimization problem. Across three LLMs, AdaqL reduces the mean bit-width by 6.75\%--10.75\% compared with uniform 4-bit RTN while maintaining comparable average performance, yielding more Pareto-efficient accuracy--bit-width operating points.