D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
Tianyu Wu ⋅ Yu Yao ⋅ Zhenting Qi ⋅ Han Zheng ⋅ Chengxi Zhang ⋅ Zhuohan Wang ⋅ Haoran Ma ⋅ Zichun Liao ⋅ Himabindu Lakkaraju ⋅ Ju Li ⋅ Yilun Du
Abstract
Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Current SOTA diffusion-based parallel drafters such as dFlash predict the full $B$-token block in one forward pass, allowing deeper drafters and higher accepted length. However, across the field, from autoregressive tree drafters to parallel block drafters, training losses use fixed per-position weights that do not adapt to the drafter's current bottleneck position. We derive per-position training weights directly from the expected accepted block length, so that each position's weight reflects its actual contribution to acceptance. The resulting loss, D-PACE (Dynamic Position-Aware Cross-Entropy), automatically identifies where the block's acceptance bottleneck lies and shifts training signal there as the drafter improves. D-PACE consistently improves both wall-clock speedup and accepted length across diverse benchmarks and positions as a drop-in replacement with negligible training overhead and requires no changes to the drafter architecture or inference pipeline, making it broadly applicable to general model architectures.
Chat is not available.
Successful Page Load