AdaTree: Serving-Aware Adaptive Tree Construction for Speculative Decoding
Abstract
Speculative decoding accelerates generative inference of large language models (LLMs) by using a small draft model to propose multiple candidate tokens, which are then verified in parallel with the target model in a single decoding iteration. While the state-of-the-art method of tree-based speculative decoding helps improve generation throughput over non-speculative inference, deploying it in LLM serving systems often yields suboptimal performance due to dynamically changing serving conditions. In particular, our analysis shows that the optimal tree configuration---the one that maximizes performance---varies with two key serving conditions: request rate and per-request characteristics. Based on the analysis, we present AdaTree, a plug-in component for LLM serving systems that dynamically adapts tree configurations to varying serving conditions. AdaTree predicts model execution time and acceptance length across different tree configurations and selects the one that maximizes speculative decoding efficiency. To capture the non-linear relationship between tree configuration and model execution time, AdaTree employs decision trees as its core modeling primitive. Our evaluation shows that AdaTree consistently outperforms both chain-based and static tree-based speculative decoding across diverse serving conditions.