Interference Beyond Geometry in Concept Extraction
Valérie Costa ⋅ Bahareh Tolooshams
Abstract
Interference is commonly treated as geometric overlap between learned features. We introduce effective interference, which combines feature geometry and code statistics to capture realized interactions. This perspective distinguishes constructive from destructive interference and frequent weak interactions from rare strong ones. Our analysis identifies four ways in which architectural constraints can shape interference: feature orthogonalization, bias compensation, norm adaptation, and encoder-decoder separation. Experiments with sparse autoencoders illustrate how constrained models selectively reduce overlap among interacting features, whereas less constrained ones exploit constructive decoder interactions for reconstruction.
Chat is not available.
Successful Page Load