Back to Basics: Testing Hessian-Based Sensitivity Scores for PoT Weight-Quantization Damage
Dominic Tarnowski
Abstract
Hessian-based curvature analyses are a tool for quantization-sensitivity ranking, typically combining curvature and perturbation magnitude into a single product, $\mathrm{Tr}(H)\cdot\|\Delta W\|^2$, to guide mixed-precision assignment. Whether it actually tracks per-layer loss damage under a specific, hardware-attractive quantization scheme is a live and practically consequential question we test directly. We isolate each layer's own weight-quantization damage by quantizing exactly that layer's weights under Power-of-Two (PoT) encoding and measuring the resulting increase in validation loss, across three CNN architectures (a compact CNN, ResNet-18, ResNet-50) on two datasets (ImageNet100, CIFAR10; post-training quantization), then test three sensitivity scores against it: raw curvature $\mathrm{Tr}(H)$, the perturbation term $\|\Delta W\|^2$ alone, and their HAWQ-V2 product. None of the three scores is universally reliable: raw curvature identifies the true single most-damaged layer at rank~1 in 3 of 6 architecture$\times$dataset combinations, and its size-normalized top-$k^*$ overlap with the truly most-damaged layers matches or exceeds the other two scores' in every combination we test; neither the perturbation term nor the HAWQ-V2 product ever exceeds it. Which specific layer is most damaged is itself dataset- and architecture-dependent. The network's stem convolution dominates on ImageNet100, but not always on CIFAR10, where damage can be small and diffuse. Full-ranking Spearman correlation is weak overall, with a few exceptions. Robust top-$k$ identification, not full-distribution correlation, is the finding that survives across datasets.
Chat is not available.
Successful Page Load