FiedlerPrune: Connectivity-Preserving Cross-Layer Pruning for Large Language Models
Abstract
Large language models (LLMs) have achieved remarkable success across diverse tasks, yet their ever-growing model sizes impose significant computational and memory costs that hinder practical deployment. Pruning addresses this by introducing sparsity into weight matrices, but existing layer-wise pruning methods operate independently on each linear layer, ignoring cross-layer dependencies and yielding globally suboptimal masks. End-to-end pruning methods address this limitation but incur prohibitive computational overhead. In this work, we propose FiedlerPrune, a pruning framework that achieves a better balance between the two paradigms by incorporating cross-layer dependencies. Guided by graph theory, FiedlerPrune constructs weighted multipartite graphs over consecutive linear layers within each Transformer block, and derives an inter-layer importance score from each weight's contribution to the Fiedler Value of the corresponding graph. Combined with any intra-layer metric, FiedlerPrune provides a unified and versatile pruning criterion reflecting both local quantitative importance and cross-layer connectivity-preserving importance . Extensive experiments demonstrate that FiedlerPrune consistently improves over layer-wise baselines, and matches or outperforms end-to-end methods at orders-of-magnitude lower pruning cost. We further demonstrate FiedlerPrune's practical utility through inference speedup, compatibility with quantization, and post-pruning fine-tuning.