LRCC: Generalizing Low-rank Compression with Conditional Computation
Abstract
Low-rank compression is a computationally inexpensive technique to reduce the number of parameters in a pretrained LLM, but uses a fixed computational structure regardless of the input activations. Conditional computation techniques address this limitation, but are often difficult to apply to pretrained networks. We introduce LRCC, a technique that trains lightweight per-block routers over a nested family of models constructed using established low-rank compression methods. We apply LRCC to Llama models and evaluate it on language modeling and several downstream tasks. At the same ratio of activated parameters, LRCC improves average accuracy by up to 7.6 p.p. on Llama-2-7B at an activated-parameter ratio of 0.9 compared to the best low-rank compression baseline, while achieving competitive predictive performance at matched batch-size-1 decoding latency.