The Price of Precision: A Calibrated Cost Model for Mixed-Precision LLM Serving
Abstract
Choosing a mixed-precision configuration for large language model (LLM) serving trades speed against accuracy. Prior allocators price that speed by a proxy such as the average bit-width or a memory budget, but served speed is governed by the binding resource (compute, weight memory, or the key-value (KV) cache), which shifts with phase, batch, and mixture-of-experts (MoE) routing, so fewer bits can be slower. We build a single cost model of served mixed-precision speed that reasons about the binding resource directly, and put it to two uses. First, as a predictor: given a configuration and operating point, it forecasts absolute served speedup. Across three hardware classes and three format families, it predicts held-out served speed within 5.0% on a 30B MoE on an AMD MI355X, within 5.5% on an MI300A, and within 9.4% on a consumer Radeon 780M, including a sign inversion (fewer bits are sometimes slower) that the bit-count proxy gets wrong. Second, we invert the same model into a per-tensor allocator: we model its per-unit accuracy and time costs as additive, and use an exact dynamic program to return the optimal configuration at a target served-speed budget or accuracy budget. Turning the decision form into absolute time adds a few constants calibrated once per GPU-and-stack.