BASTION: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting
Soowon Oh ⋅ Nam Cao ⋅ Yujin Kim ⋅ Hojung Jung ⋅ Huzama Ahmad ⋅ Sangmin Bae ⋅ Se-Young Yun
Abstract
Block-diffusion drafters have recently shown strong potential for speculative decoding by predicting multiple future-token distributions in a single forward pass. To exploit these parallel predictions, tree-based verification can check multiple plausible continuations instead of committing to a single greedy draft path, often increasing the number of accepted tokens per decoding cycle. However, existing tree-based approaches typically rely on fixed tree topologies or static verification budgets, even though the best budget depends on the local draft distribution, context length, target model, and hardware runtime. We propose a system-aware adaptive tree framework for block-diffusion speculative decoding. At each decoding cycle, our method builds a tree toward high-probability draft paths and selects the verification budget that maximizes an online speedup estimate. This estimate combines a drafter-based acceptance surrogate with an analytical verifier-latency model calibrated from observed runtimes. The framework requires no additional training of the drafter or target model and preserves the target model's decoding rule. Across target models, benchmarks, decoding temperatures, and GPU platforms, our method achieves up to a $6.93\times$ wall-clock speedup over standard autoregressive decoding. This corresponds to approximately 40% higher speedup than Dflash, the strongest existing block-diffusion drafter, and is obtained without per-setting budget tuning.
Chat is not available.
Successful Page Load