MixNLQ: An Effective Post-Training Nonlinear Low-Bit Quantization Method for Large Language Models
Abstract
Post-Training Quantization (PTQ) efficiently compresses Large Language Models (LLMs) without expensive retraining. However, most existing PTQ methods rely on linear quantization, which fails to fully capture the underlying characteristics of parameter distributions. Conversely, non-linear approaches, such as codebook-based methods, introduce significant complexity that hinders practical engineering deployment. To address this, we propose MixNLQ, an easy-to-use hybrid non-linear PTQ method. By combining linear and exponential terms, MixNLQ flexibly fits the "high-peak, heavy-tail" distributions typical of LLM weights and activations. It supports both calibration-free and calibration-based settings, and uses arithmetic-only dequantization to avoid memory-intensive lookup tables—ensuring highly efficient quantization and inference. Additionally, MixNLQ serves as a plug-and-play module capable of enhancing various existing PTQ frameworks. Evaluations at 3-bit and 4-bit on LLaMA, Mistral, and DeepSeek demonstrate that MixNLQ achieves competitive or superior perplexity and zero-shot accuracy compared to existing PTQ baselines, alongside promising activation quantization performance and its versatility as a plug-and-play enhancement module. Ultimately, MixNLQ offers a highly practical solution that effectively balances accuracy, algorithmic simplicity, and real-world deployment efficiency.