Learning Structured Sparsity Beyond SALAAD: Bridging Model Compressibility and Efficient Inference
Abstract
Scaling laws continue to reward increases in model scale, while the resulting computational and memory demands drive greater reliance on cloud infrastructure for inference. Yet such reliance becomes impractical under real-world constraints on memory, latency, and privacy, motivating local execution on resource-constrained devices. To alleviate memory bottlenecks, an emerging line of work learns sparse-plus-low-rank representations of model weights during training; SALAAD, in particular, enables elastic memory--perplexity trade-offs under subsequent compression. However, fewer parameters do not necessarily lead to faster inference, as the irregular memory accesses and incompatibility with structured accelerator primitives induced by unstructured sparsity can erase the expected computational savings and even slow execution. To bridge this gap, we introduce S-SALAAD, which jointly induces block-sparse and low-rank components during training. We provide a practical implementation of S-SALAAD across training, elastic compression, and inference, with a dedicated GPU kernel enabling block-sparse acceleration. Across LLaMA-1B pretraining and Qwen3-1.7B cross-architecture distillation, S-SALAAD establishes superior perplexity--memory--throughput Pareto frontiers over existing baselines. Together, these results show that S-SALAAD overcomes the practical limitations of unstructured sparsity, bridging model compressibility and efficient inference.