Iterative ILP with Update-Size Control for Reducing Surrogate-Task Mismatch in Bit-Width Selection
Abstract
Mixed-precision quantization improves the accuracy-efficiency trade-off by assigning different bit-widths across a network. Since bit-width selection is a combinatorial problem, existing methods often optimize additive surrogate losses instead of directly evaluating task losses in all configurations. We show that these surrogates can suffer from surrogate-task mismatch, especially when many layers are changed simultaneously from a reference configuration. We further observe that, although allowing more layers to change improves the surrogate objective, the evaluated task loss can be minimized at an intermediate limit on the number of changed layers. Based on this observation, we propose Iterative Integer Linear Programming (I2LP), which iteratively refines a reference bit configuration. At each refinement step, I2LP solves multiple ILPs with different limits on how many layers may change from the current reference, and accepts the best candidate only when it reduces the task loss. I2LP applies to both post-training quantization and quantization-aware training, and consistently outperforms uniform-precision baselines and existing mixed-precision methods. Code will be released.