How Does Pruning Change Decisions in Large Language Models?
Abstract
Post-training pruning reduces the memory and computation costs of large language models (LLMs), and N:M semi-structured sparsity provides a hardware-friendly pruning pattern for efficient inference. However, pruned models are often evaluated mainly by perplexity and average downstream accuracy, which are useful aggregate metrics but do not fully explain how pruning changes individual candidate-level decisions. We evaluate magnitude pruning, Wanda, SparseGPT, Wanda++, and our proposed MR-GOBS across multiple model families and sparsity settings. We further conduct a paired decision-level evaluation that measures true-label confidence, answer margin, prediction flips, candidate-level distribution shift, and near-boundary vulnerability. Our results show that pruning weakens decision reliability by lowering confidence in the correct answer, shrinking answer margins, and increasing correct-to-wrong flips, especially on near-boundary examples. To examine whether reconstruction-oriented improvements reliably improve decision reliability, we revisit SparseGPT through its local least-squares reconstruction formulation and refine its N:M mask-selection step. MR-GOBS uses GroupOBS, a set-level OBS reconstruction cost for jointly pruning a group of weights, to score each legal local N:M mask. It then selects masks using a minimax-regret criterion over multiple calibration views. Although MR-GOBS often improves average perplexity over SparseGPT, these gains do not always translate into better downstream accuracy or decision-level metrics. Our findings suggest that reliable evaluation of pruned LLMs requires both aggregate performance metrics and decision-level diagnostics.