Scattered by Design: Why Per-Output Pruning Resists Structured Compression
Abstract
Unstructured pruning methods such as Wanda preserve model quality but produce irregular sparsity with no practical compression, while structured methods enable efficient inference at the cost of degraded performance. We ask whether post-hoc linear transformations can convert unstructured sparsity into structured patterns, combining both benefits. Across a broad class of fixed linear transformations—including rotations, decompositions, and permutations—we find no increase in structured sparsity (Δ = 0.0%) in any OPT feed-forward layer. Transformations that yield structure do so identically on dense matrices, independent of pruning. We attribute this to three factors: rank preservation, weak inter-row mask correlation, and spectral dispersion. Together, these prevent post-hoc transformations from inducing structure, suggesting that compressible sparsity must be introduced during pruning rather than recovered afterward.