A Curvature Phase Transition Governs Coherence Penalty Efficiency Against Feature Absorption in SAEs
Hak Hyun Kim ⋅ Yash Raj ⋅ Peter Chin ⋅ Soroush Vosoughi
Abstract
Feature absorption is a structural failure mode of sparse autoencoders, where a single dictionary atom captures multiple distinct concepts rather than one, reducing interpretability. While prior work has shown that absorption configurations are spurious local minima of the sparse dictionary learning loss, little is known about the geometric mechanisms that determine whether a coherence penalty can escape them. We study the curvature geometry of pairwise coherence penalties at absorption configurations on the unit sphere. We show that (1) every coherence penalty in the power family $\mathcal{R}p = \sum{l 0$, yet the efficiency of this force depends critically on the penalty's curvature; (2) a phase transition at $p = 2$ governs this efficiency: the Riemannian Hessian in the splitting direction equals $p(p-2) \cdot k^{1-p/2}$ in closed form, so concave penalties ($p < 2$) receive curvature assistance and convert absorption into a strict saddle point, while convex penalties ($p > 2$) face curvature resistance and deepen the local minimum; (3) the boundary at $p = 2$ corresponds precisely to a Parseval invariance that renders the quadratic penalty geometrically neutral; and (4) under feature correlation $\mu^{}$, convex penalties reverse above $\mu^{}_{crit} = (p-2)/(p+k-2)$, giving a closed-form threshold for when resistance can be overcome. Experiments on Pythia-160M and Gemma-2-2B confirm the predicted $\lambda$-efficiency ordering for competitive-activation SAEs, with absorption reduction verified as feature-level improvement rather than feature suppression. More broadly, our results point toward a principled framework for selecting coherence penalties based on the curvature geometry of the loss landscape rather than empirical tuning.
Chat is not available.
Successful Page Load