Learning Cost-Efficient Autoscaling for Latency-Constrained Disaggregated LLM Serving
Abstract
Prefill-Decode (PD) disaggregation improves LLM serving by isolating compute-bound prefilling and memory-bound decoding onto specialized GPU pools. However, dynamically rebalancing these pools is challenging because request arrival rates, prompt lengths, and generation lengths change over time. We formulate this autoscaling problem as latency-constrained cost minimization: the controller must minimize GPU usage while keeping P99 tail latency within TTFT and TPOT targets. Existing threshold-based heuristics require extensive parameter tuning, yet still tend to over-provision. To address this, we propose a reinforcement learning (RL) framework that learns a cost-efficient autoscaling policy. The policy uses a recurrent network to track multi-timescale workload trends and pod lifecycle states, enabling decisions that account for delayed pod readiness and delayed SLO feedback. It also restricts the action space with hard masks and anti-oscillation cooldowns, enforcing feasible scaling actions while reducing wasteful GPU transitions. We train the policy in a high-fidelity PD-disaggregated serving simulator calibrated from profiling data and validated against a physical testbed. Simulator experiments on ShareGPT and Azure Conversational traces show that the learned policy satisfies the same latency constraints as grid-searched heuristic baselines while reducing average GPU occupancy. Real-cluster deployment on ShareGPT further confirms sim-to-real transfer while satisfying the latency targets.