SLORR: Simple and Efficient In-Training Low-Rank Regularization
Abstract
Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. Existing training-time low-rank regularizers can improve compressibility, but they often require SVDs of large weight matrices, modify the model architecture (introducing additional trainable parameters), or rely on stateful cached quantities. To address these limitations, we introduce SLORR, a simple, stateless, and architecture-preserving framework for in-training low-rank regularization, instantiated with two main variants based on the Hoyer sparsity metric and the nuclear norm. SLORR directly regularizes the original weight matrices using GPU-friendly approximations for the forward and backward passes of the regularizers, for which we provide approximation guarantees. We first evaluate SLORR on ImageNet-1K across short-horizon continued training and pretraining of ResNets and Vision Transformers, where SLORR induces compressibility while introducing less than 8\% training overhead. We further evaluate SLORR on LLM pretraining at 135M, 560M and 1B scales: SLORR-trained compressed models preserve performance substantially better than unregularized models while adding less than 3\% average training overhead.