GASP: Group-regularized Adaptive Structured Pruning
Leandro Palma ⋅ Kyle Poe ⋅ Salma Tarmoun ⋅ Lachlan MacDonald ⋅ Benjamin D Haeffele ⋅ Rene Vidal
Abstract
State-of-the-art structured pruning methods often rely on greedy heuristics and require careful tuning of multiple hyperparameters. In contrast, regularization-based methods introduce sparsity-inducing regularization to the training objective, naturally promoting group sparsity in the network parameters. This principled strategy enjoys convergence guarantees, but requires a careful design of the pruning groups in the regularizer and is extremely brittle to the exact value of regularization strength, drastically limiting its adoption. In this work, we propose \textit{Group-regularized Adaptive Structured Pruning (\texttt{GASP})}, a principled pruning method that challenges the supremacy of costly, state-of-the-art heuristics by demonstrating that group-sparse regularization achieves competitive performance when (1) the pruning groups are chosen to account for the invariance in the scaling of parameters of the network architecture and (2) the regularization strength is dynamically adapted to the $\ell_2$ norms of the groups, reducing costly hyperparameter tuning. Moreover, we demonstrate the generality and robustness of \texttt{GASP} across model families, scales, tasks, datasets and architectures by pruning well-established pretrained models in both language and vision domains, including reading comprehension on SQuAD, reasoning benchmarks such as PIQA, HellaSwag, WinoGrande, ARC-e and ARC-c, human pose estimation on COCO, and image classification on ImageNet-1K.
Chat is not available.
Successful Page Load