Decision-Aware Rounding for LLM Quantization
Abstract
Post-training quantization (PTQ) reduces the memory and computational cost of large language models, but most PTQ objectives optimize local numerical distortion rather than downstream task decisions. We propose Decision-Aware Rounding (DAR), a task-aware adaptive rounding objective that projects weight-space quantization perturbations onto full-precision pairwise answer-margin gradients and penalizes margin-reducing movement relative to the available FP margin. Using 3-bit group-wise weight quantization on the final two transformer blocks, DAR reduces harmful FP-to-quantized decision flips on Qwen2.5-0.5B-Instruct with ARC-Challenge from 26.72% to 19.31% across three calibration seeds while maintaining mean task accuracy (42.92% vs. 43.14%). On SmolLM2-360M-Instruct, harmful flips decrease from 64.65% to 39.39%, while accuracy increases from 24.97% to 28.87%, with both metrics improving in all three seeds. We also observe smaller positive transfer on HellaSwag. Finally, held-out first-order projections strongly predict actual post-quantization margin changes, supporting FP answer-margin geometry as a measurable task-relevant signal for low-bit rounding.