Coverage-Based Calibration for Post-Training Quantization via Weighted Maximum Coverage over Outlier Channels
Ibne Farabi Shihab ⋅ Sanjeda Akter ⋅ Anuj Sharma
Abstract
Post-training quantization (PTQ) compresses large language models to low bit-widths using a small calibration set, making calibration selection a consequential but under-specified part of the quantization pipeline. We study calibration selection for weight-only LLM PTQ backends such as AWQ and GPTQ. In this setting, calibration data does not define an inference-time activation clipping threshold; instead, it determines activation-conditioned quantities used for weight reconstruction, saliency estimation, and scale search. We identify a failure mode in which a calibration set misses rare, high-magnitude input channels of quantized linear modules, causing the backend to underrepresent reconstruction-sensitive directions. Motivated by this observation, we formulate calibration selection as weighted maximum coverage over module-level activation outlier channels. Each candidate sequence covers the module/channel pairs it activates above a reference threshold, and each covered pair is weighted by a dimensionally consistent surrogate for its potential contribution to squared output reconstruction error. The resulting objective is monotone submodular, so greedy selection admits the classical $(1-1/e)$ approximation guarantee for the coverage objective. We instantiate this formulation in COVERCAL, a backend-agnostic calibration selector operating on cached activation masks. Across LLaMA-2, LLaMA-3, and Mistral models, under AWQ and GPTQ INT4 backends and five downstream evaluations, COVERCAL improves over the reported fixed-pool calibration baselines, with the largest gains at small calibration budgets. Under AWQ at $K=128$, COVERCAL improves MMLU by $1.2$--$1.5$ points over random calibration; at $K=64$, it matches or exceeds random calibration at $K=256$ on LLaMA-3-8B. The contribution is not a new quantizer, but a module-correct coverage formulation, an efficient greedy selector, and empirical evidence that weighted outlier coverage is a useful calibration signal for existing weight-only PTQ backends.
Chat is not available.
Successful Page Load