CORP: Closed-Form One-shot Representation-Preserving Structured Pruning for Transformers
Abstract
Transformers achieve strong accuracy but incur high compute and memory cost. Structured pruning reduces inference cost, but most methods rely on retraining or multi-stage optimization, which limits post-training deployment. We propose CORP, a closed-form one-shot structured pruning method that removes MLP dimensions and attention substructures using only unlabeled calibration data without labels, gradients, or fine-tuning. CORP formulates structured pruning as a representation recovery problem. It models removed components as affine functions of retained components and derives closed-form ridge regression solutions that fold compensation into model weights. This minimizes the expected representation error under the calibration distribution. Experiments on ImageNet with DeiT reveal strong redundancy in both MLP and attention representations. Without compensation, one-shot structured pruning causes severe accuracy loss. With CORP, models retain high accuracy under aggressive sparsity. On DeiT-Huge, CORP achieves 83.27% Top-1 accuracy after pruning 50% of both MLP and attention structures.