The Sparsity Whisperer
Abstract
Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also pairwise output differences. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across the Llama 2 and 3.1 families, our second-order variant consistently improves language modeling performance over strong reconstruction-based baselines, while our update-free variants improve over activation-aware update-free baselines, with gains stronger in more constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families, suggesting that preserving input differentiation is a broadly useful signal for post-training LLM sparsification.