Lightweight Quantization Rotations with a Distortion Transfer Bound
Chandan Sah ⋅ ARPAN VERMA ⋅ Dev Utkarsh Pal ⋅ ADITYA UPADHYAY ⋅ Harshit Gupta
Abstract
Small language models still spend most of decode memory on the KV cache, and most of the rotation cost of a ``near-optimal'' scalar quantizer on a dense Haar matrix. We study the two-stage Randomized Hadamard Transform as a lightweight, data-oblivious replacement: $2d$ sign bits, $O(d\log d)$ arithmetic, no calibration set. For symmetric nearest-centroid squared error the distortion gap to Haar is $O(1/d)$ on generic inputs, not the $O(d^{-1/2})$ Berry--Esseen rate of the coordinate laws. We prove this at $1$ bit from an exact moment cancellation, and we measure $1/d$ scaling, GPU latency, and $2$--$4$-bit grouped KV on Qwen2.5-0.5B and TinyLlama-1.1B. At $d=1024$ the rotation state is $0.25$ KB versus $8$ MB for a dense float64 matrix. The trust claim is narrow: the map is data-oblivious and comes with an explicit transfer bound versus Haar-optimal scalar quantization. It is not a certificate for $2$-bit caches, for weight compression, or for calibrated/learned rotations.
Chat is not available.
Successful Page Load