LoopPrune: Evolutionary Module Selection with Iterative Execution for Efficient Large Language Model Compression
Abstract
Structured pruning offers hardware-agnostic acceleration for Large Language Models (LLMs) but typically degrades performance by strictly adhering to the original static sequential execution graph. In this paper, we challenge this rigidity with LoopPrune, an evolutionary framework that transforms model compression into an architectural re-discovery process. By treating pre-trained layers as a library of reusable primitives, LoopPrune optimizes a dynamic execution path where modules can be skipped, executed once, or iterated via controlled loops. This approach effectively decouples effective model depth from parameter count, allowing compressed models to maintain high reasoning capacity. To navigate the vast combinatorial search space, we introduce a hybrid evolutionary strategy enhanced by a Multi-Tier Population Initialization mechanism, which stratifies the search start-point using cross-model elite transfer, structural priors, and Wanda score importance estimation. Extensive experiments on diverse LLM families demonstrate that LoopPrune sets a new state-of-the-art across all sparsity levels. Notably, LoopPrune identifies efficient iterative patterns that allow compressed models to match or even surpass the zero-shot reasoning accuracy of their dense counterparts on specific benchmarks, while maintaining near-lossless perplexity at high retention rates.