Generation-Drift-Guided Block Pruning for Large Language Models
Abstract
Pruning reduces the model size and inference cost of large language models (LLMs), yet preserving performance across diverse downstream tasks remains challenging. Existing structured pruning methods commonly rely on local discrepancy measures or teacher-forced predictive criteria, such as Block Influence (BI) and perplexity (PPL). While these methods are effective on classification benchmarks, the resulting pruned models often degrade severely on generative reasoning tasks. We study redundancy in Transformer-based models and observe that pruning-induced generation drift is strongly correlated with the functional importance of modules, suggesting a more faithful criterion for identifying redundant modules. Based on this insight, we propose GDPruner, a calibration-free, generation-drift-guided block pruning method. GDPruner constructs lightweight self-generated probe trajectories, estimates module importance using tail-position generation drift, and removes redundant modules through adaptive search. Extensive experiments show that GDPruner surpasses state-of-the-art structured pruning baselines while better preserving overall performance and complex generative reasoning ability, offering a promising direction for generation-robust LLM deployment.