TopK Sparsity Costs Predictive Power and Overfits Its Own Selection Procedure at Low K
Abstract
Sparse autoencoders (SAEs) are increasingly used to decompose biological foundation models into interpretable features. This carries an implicit assumption: since reconstruction fidelity can be pushed arbitrarily high, a sparse representation should be a safe stand-in for the dense representation it was trained on. We test this assumption in a concrete, practically motivated drug-discovery setting---filtering poor docking poses from a physics-informed protein-ligand docking pipeline's variational autoencoder latent space, across a sweep of sparsity levels---and find that it does not hold. Dense latents consistently match or outperform sparse autoencoder representations on pose-quality classification, even at sparsity levels where reconstruction is nearly perfect. Separately, pose filters built from the sparsest, and by the usual argument most interpretable, representations pass their own selection procedure but systematically fail to generalize to held-out data, while less aggressively sparse representations hold up reliably. Both point to a shared explanation: TopK's magnitude-based feature selection discards classifier-relevant information and leaves individual features with too few examples for reliable selection-set statistics. The practical caution for anyone deploying SAE-based filters in biological pipelines is direct: reconstruction fidelity and feature-level interpretability are not proxies for downstream predictive power or selection reliability, and both must be checked directly, on the actual probe and selection procedure in use, at every sparsity level under consideration, before deployment.