StallBound: The Memory-Bandwidth Limit of Chunked Prefill
Abstract
Prefilling a new prompt can delay the next token of other requests that share the device. Chunked prefill bounds this delay by splitting the prompt into small chunks and running decode between them. Each chunk step reads the full model weights, so chunk size cannot push the delay below the time to read those weights. We measure the best achievable time between tokens (TBT), the smallest worst-case decode gap over chunk sizes, for dense models of 1 to 28 GB on a unified-memory system and a discrete GPU. Best TBT is linear in the weight bytes read per step, and its slope tracks memory bandwidth: the two devices' slopes differ by 5.76× against a 5.86× bandwidth ratio. Decode-only fits recover 96–99% of measured bandwidth and predict nine additional dense, quantized, other-family, and mixture-of-experts configurations within 8%. Solving the fitted line for a TBT target gives the largest model a device can serve; at 50 ms, that limit differs by 6.1× between the two devices. A model beyond it misses the target at every chunk size and needs fewer weight bytes, a shorter prompt, or more bandwidth.