Output-Aware Block Influence: Jacobian-Lens Weighting for Depth Pruning
Kyle Lemoi ⋅ Jonathan Yong ⋅ Shubhankar Tripathy ⋅ Priestley Fernandes ⋅ Anirban Majumder ⋅ Zarreen Reza
Abstract
ShortGPT is a fundamental pruning approach for removing layers. It defines a score, Block Influence (BI), measuring the hidden state rotation across a layer. While effective, it is a local measure and does not account for downstream impacts. We address this limitation using the Jacobian Lens (J-Lens), a map transporting perturbations at a layer to the final residual stream. At a common sparsity level, pruning scores applying J-Lens transportation to $BI$, $J-BI$, were found to improve post-pruning accuracy up to 19 points over $BI$ while reducing damage to perplexity for Qwen3-8B, Llama-3.1-8B-IT, and Gemma-3-12B-IT. We then perform a single-block ablation study and use J-Lens probing to propose an explain why the proposed scores improved performance.
Chat is not available.
Successful Page Load