Dense Structural Compression of Transformers via Gauge-Correct Channel Removal
Jed A Duersch ⋅ Naïm Es-sebbani ⋅ Nathanaël Haas ⋅ Zied Bouraoui
Abstract
We develop a methodology from first principles to adapt transformer structural complexity during training and maximize inference utility per unit compute. Calibrated channel-level penalties compel operators to reorganize into smaller dense tensors, driving entire tensor slices to zero to enable physical removal of channels while preserving density and full GPU throughput. The natural approach, penalizing the norm of operator components that act through each channel, is provably destabilized by gauge freedom between multiplicatively coupled factors. We resolve this pathology with additive symmetric group-lasso penalties that recover a monotone function of product-norms when the network converges to gauge balance. Our equilibrium analysis shows how to calibrate per-channel penalty strength to correctly suppress channels that under-perform in inference utility per unit compute. Under adaptive penalty pressure, the network reorganizes its representations into depth-dependent structural profiles that can be far smaller than the architecture required to learn the task. On polynomial long division over~$\mathbb{F}_{31}$, the method achieves $190{\times}$ compression with perfect accuracy. On character-level language modeling and masked autoencoding, compressed models match or exceed hand-designed baselines at equal inference cost. Continuous dense compaction accelerates training itself, with step times decreasing as the model compresses. Post-hoc pruning with the same utility metric cannot reach these architectures, confirming that the compressed representations emerge from sustained resource pressure, not from removal of redundancy from a fixed solution.
Chat is not available.
Successful Page Load