Accelerating Long-Context LLM Prefill via Layer-wise Progressive Token Pruning in Local Deployment
Abstract
Long-context inference is increasingly important for privacy-sensitive local LLM deployment. In this setting, the prefill stage dominates Time-to-First-Token (TTFT), yet existing acceleration methods struggle to balance speed and quality. Sparse-attention methods accelerate only the attention module, leaving Feed-Forward Network (FFN) costs unchanged, while auxiliary-model-based compression often compromises quality under tight memory budgets. Recent layer-wise token pruning methods reduce both attention and FFN computation, but they typically rely on heuristic pruning schedules or token-recovery mechanisms that introduce decoding overhead. We propose FastPrefill, a recovery-free layer-wise token pruning system for efficient local long-prompt inference. FastPrefill employs an offline data-driven optimizer to determine layer-wise pruning schedules that maximize fidelity under a latency constraint, and executes these schedules via an algorithm–system co-design supporting head-specific sparse attention at runtime, where each head dynamically attends to its own selected key/value tokens. Evaluations on LongBench and RULER show that FastPrefill maintains comparable inference quality while achieving up to 2.13x TTFT speedup and 3.03x end-to-end speedup over state-of-the-art baselines.