Zero-Shot Quantization via Weight-Space Arithmetic
Daniele Solombrino ⋅ Antonio Andrea Gargiulo ⋅ Alessandro Zirilli ⋅ Luca Zhou ⋅ Robert Adrian Minut ⋅ Emanuele Rodolà
Abstract
We show that robustness to post-training quantization (PTQ) is a transferable direction in weight space. We call this direction the \emph{quantization vector}: extracted from a donor task by simple weight-space arithmetic, it can be used to patch a receiver model and improve post-PTQ Top-1 accuracy by up to \(60\)$\%$ in a 3-bit setting, without receiver-side quantization-aware training (QAT). Because the method requires no receiver training data, it provides a zero-shot, low-cost alternative to QAT for extremely low-bit deployment. Across multiple vision/language models and more than 30 tasks, donor quantization vectors often yield substantial gains even when donor and receiver tasks differ markedly. We further prove rigorously that quantization vectors are well-defined and do not suffer from reparameterization symmetries, and provide a local geometric account of their effects. Together, these results suggest that quantization robustness can be partially isolated, reused, and transferred through simple weight-space algebra.
Chat is not available.
Successful Page Load