Keep or Preempt? Termination-Aware Scheduling for LLM Serving with Speculative Decoding
Ziyu Cheng ⋅ Songtao Guo ⋅ Mingyan Li ⋅ Chang Han ⋅ Hongyu Xu ⋅ Xu Luo
Abstract
Effective inference is critical for interactive large language model (LLM) serving, where both time-to-first-token (TTFT) and end-to-end (E2E) latency shape user experience. Speculative decoding has emerged as a promising solution that reduces inference latency by employing lightweight drafting and parallel verification. However, it introduces a new scheduling issue, i.e., traditional shortest remaining processing time (SRPT) scheduling based on output length estimation fails to work due to its paradigm shift in inference structure. In this paper, we explore the preemption mechanism of speculative decoding-based LLM serving systems and show that — besides the output length, a request's remaining service time depends on draft acceptance and the number of verification rounds. Building on this insight, we develop a novel termination-aware scheduling framework, named \textsc{AlmostDone}, which first \textit{identifies requests that are almost done} within several future local decoding steps. It achieves this by introducing a lightweight predictor built from runtime draft-and-verify states, which helps approximate the SRPT-like scheduling better than existing approaches. Then, \textsc{AlmostDone} \textit{determines whether preemption is worthwhile} through a switching-aware online policy to limit excessive running-set alterations. Experiments on an open-source vLLM serving system show \textsc{AlmostDone} substantially surpasses the default schedulers, reducing mean and median TTFT by up to 3.83$\times$ and 3.39$\times$, and mean and median E2E request latency by up to 1.39$\times$ and 2.33$\times$. Against the oracle length-based preemption baseline TRAIL+, it still achieves up to $2.0\times$ lower mean TTFT and $1.3\times$ lower mean E2E request latency.
Chat is not available.
Successful Page Load