Block-OBS-GS: Exact Per-Block Joint Brain Surgery with Gauss–Seidel Refinement for LLM Pruning
Yuwen Huang ⋅ Xiang Pan
Abstract
Post-training pruning of large language models (LLMs) zeroes a fraction of the weights and reconstructs the survivors without retraining, addressing the weight-memory bottleneck of LLM serving. At a fixed mask the per-row reconstruction admits a closed-form Optimal Brain Surgeon (OBS) optimum, but the cost of the exact joint solve is cubic in the layer dimension; one-pass methods such as SparseGPT and Wanda approximate it at quadratic cost without attaining the per-block joint optimum. We introduce Block-OBS-GS, which solves the per-block joint OBS reconstruction exactly within SparseGPT's complexity class via a per-row Cholesky factorisation of the in-block damped Hessian (one third the cost of an explicit inverse, reused across the algorithm) and follows it with a single cross-block Gauss--Seidel sweep. We prove three guarantees: per-layer FLOP count in SparseGPT's complexity class, with the leading $mn^{2}$ constant dropping from $1/2$ (SparseGPT) to $3\rho/2$ (Block-OBS-GS) where $\rho=1-s$, hence strictly lower for sparsity $s>2/3$; per-row reconstruction error at most SparseGPT's, strict whenever the activation Hessian has non-zero off-diagonal coupling; and a closed-form distance bound to the unregularised joint-OBS optimum that contracts geometrically per sweep. Across six Llama-3 and Qwen-3 backbones from 3B to 70B parameters, Block-OBS-GS attains the lowest WikiText-2 perplexity at every tested sparsity on the three Llama-3 backbones, with a $47\%$ perplexity reduction over SparseGPT on Llama-3.1-8B at $s=0.875$; the highest seven-task commonsense Avg7 on six of twelve populated cells; 5-shot Massive Multitask Language Understanding (MMLU) within $0.026$ of SparseGPT across populated cells; and the lowest WikiText-2 perplexity on four of five backbones under $2{:}4$ structured sparsity.
Chat is not available.
Successful Page Load