COALA 2: Adaptive Rank Selection for Context-Aware Low-Rank Approximation
Abstract
Context-aware low-rank approximation has recently attracted considerable attention as a technique for compressing large neural networks. Although recent methods provide numerically stable solutions for constructing such approximations, they typically assign a separate rank budget to each layer, despite the fact that GPU memory is constrained at the level of the entire model. Existing methods for non-uniform rank allocation also face important limitations, including inaccurate approximations and dependence on proxy objectives that may not directly reflect the final model quality. We address these challenges with an efficient framework for adaptive rank selection that augments state-of-the-art low-rank compression methods with trainable layer-wise ranks under a global memory budget. By learning layer-wise ranks directly from the model loss, our method adapts the rank allocation to model level performance under a global compression budget. Experiments across multiple model families, scales, and datasets show strong results in LLM weight compression and improved perplexity when the method is applied to KV-cache compression. The source code is available at: \url{https://anonymous.4open.science/r/coala2-F545}.