Grand Challenge: Predict Where the Next Bit Should Go, Before You Quantize
Jundong Hu ⋅ Shekar Ramachandran
Abstract
Post-training quantization creates a resource-allocation decision for each deployment: given a small precision budget above a low-bit (e.g. 4-bit) baseline, how should that additional precision be allocated: spent globally on finer quantization granularity throughout the model (e.g. group-128 scales), or locally on restoring a few high-value layers to higher precision? Current choices rely either on an expensive per-layer quantize-and-evaluate sweep or on a heuristic. The goal is to predict, from a pre-quantization model, (i) whether its quantization damage is diffuse or concentrated and (ii), if concentrated, which layers to restore, using architecture and cheap forward-pass statistics without running a quantize-and-evaluate sweep.
Chat is not available.
Successful Page Load