HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
Abstract
Quantization-aware training (QAT) is a key approach for adapting large language models (LLMs) to extremely low-bit weights, where post-training quantization often suffers severe accuracy loss. However, most extremely low-bit QAT methods rely on the straight-through estimator (STE), which keeps hard round-and-clip weights in the forward pass while using an identity surrogate in the backward pass, creating a mismatch between discrete forward states and continuous latent-weight updates. As a result, small updates may fail to cross quantization boundaries, causing dead-zone stagnation in sub-optimal discrete states. To address this problem, we propose Hestia, an operator-faithful differentiable QAT framework for extremely low-bit LLMs. Instead of repairing the gradient of a hard quantizer, Hestia relaxes the quantizer itself as a temperature-controlled Softmax expectation over the same discrete codebook. The relaxation provides forward--backward consistent gradients during training and anneals to the target hard quantizer for inference. To stabilize this soft-to-hard transition, Hestia further uses offline Hessian traces as lightweight tensor-wise sensitivity signals for adaptive temperature scheduling. Experiments on Llama-3.2 show that Hestia consistently improves over ternary QAT baselines, with relative average zero-shot gains of 5.39% and 4.34% on 1B and 3B models, while adding only about 0.2% training-time overhead and no inference-time overhead. The code is available at https://github.com/hestia2026/Hestia.