PerQ: Inverse Generative Modeling for Neural Image Compression via Quantization Error Compensation
Abstract
Neural image codecs face a long-standing trade-off: distortion-optimized designs preserve pixel-wise fidelity but tend to produce blurred reconstructions at low bitrates, while generative designs close the perceptual gap with a learned prior, but with pixel fidelity saturating as bitrate increases. The latter typically requires training multiple models for various bitrate targets. This paper addresses these limitations by reframing the perceptual reconstruction problem. Starting from the quantization bound that holds for any modern multi-rate codec, we derive a transport bound that confines the perceptual reconstruction to a closed, rate-determined cell around the backbone's output. The cell is small relative to the full image manifold, casting unconstrained generative modeling as a bounded-support problem. We instantiate this as PerQ: an image codec that attaches a rate-conditioned flow-matching compensator supported on this cell to a frozen multi-rate distortion-optimized backbone. The experimental results across Kodak and CLIC2020-test datasets show that PerQ offers similar performance to generative neural image codecs (e.g., MS-ILLM) in perceptual metrics (LPIPS, FID) , while still maintaining competitive coding performance compared to distortion-optimized codecs (e.g., ELIC). Moreover, PerQ can produce compression results across a wide bitrate range and fast decoding from only a single checkpoint.