The $1/\mathcal{W}$ Law: Context Length is the Dominant Energy Lever in LLM Inference Fleets
Huamin Chen ⋅ Xunzhuo Liu ⋅ Yuhan Liu ⋅ Junchen Jiang ⋅ Steve Liu ⋅ Chenxu Niu ⋅ Bowei He
Abstract
Energy efficiency in LLM inference is often treated as a property of the model, the hardware, or the serving software. We show that, for decode-dominated serving, a simpler systems variable dominates: the serving context window. Because KV-cache memory is finite, increasing the context window reduces the number of concurrent sequences a GPU can hold. Under continuous batching, this concurrency limit directly sets decode throughput, while GPU power remains nearly flat over most of the operating range. The result is a simple $1/\mathcal{W}$ law: tokens per watt scale inversely with the serving context window $\mathcal{W}$. We derive this law from KV-cache capacity and a logistic GPU power model, and empirically verify it for Llama-3.1-70B inference over the 2K--128K context range. On identical H100 hardware and software, tok/W varies by nearly $40\times$ across this range. We then show that the same mechanism determines GPU fleet-level energy efficiency. Context-length routing keeps short requests on high-concurrency short-window pools, while newer hardware shifts the power and memory curve upward. These two operator-controllable levers are approximately multiplicative: on the short-context-dominant Azure trace, two-pool context-length routing improves tok/W by $2.5\times$ over a homogeneous NVIDIA H100 fleet, a calibrated NVIDIA B200 deployment improves tok/W by $2.0\times$ at fixed topology, and combining them reaches $5.76\times$. Finally, we analyze active-parameter weight streaming in MoE models as a third, architecture-level lever, giving an upper-bound gain of up to $5.1\times$. Together, these results suggest a deployment order for energy-efficient LLM serving: route by context length first, upgrade hardware second, and exploit sparse architectures where dispatch costs are controlled. We release code, data, and simulator in anonymous [repository](https://anonymous.4open.science/r/gpu-fleet-sim-1BEB) to facilitate reproduction and further research.
Chat is not available.
Successful Page Load