ASQ: Agent-guided Semantic-aware Quantization for Large Language Models
Abstract
Post-training quantization (PTQ) compresses large language models (LLMs) but remains largely semantic-agnostic, relying on global activation magnitude to guide precision allocation. This overlooks the functional heterogeneity of Transformer representations, where semantically meaningful channels are sparsely activated while structurally dominant channels exhibit large but context-insensitive activations. We propose Agent-guided Semantic-aware Quantization (ASQ), a PTQ framework that conditions on semantically informative positions identified by an auxiliary LLM. A dual-pathway scoring mechanism captures both activation selectivity and attention routing toward these positions, producing channel-wise importance scores for scaling. This yields a closer approximation to a semantically conditioned objective and consistently improves low-bit (e.g., W4A16) quantization performance across LLMs. Extensive experiments show consistent gains over state-of-the-art PTQ, preserving general reasoning and even surpassing FP16 performance on several benchmarks in specialized domains, while enabling training-free, privacy-preserving self-quantization with zero inference overhead.