Autointerp Simulation Scores Track Token Naming, Not Explanation Truth
Abstract
Autointerp simulation scoring is used as a quality signal for sparse autoencoder latent descriptions, a use that presumes the score tracks whether a description is correct. We test that presumption with a within-latent 2x2 over whether a description names the latent's trigger token and whether its semantic gloss is true, on 16 token-driven latents of a Gemma-2-2B layer-12 GemmaScope SAE. Naming the token raises the score by 0.305, and does so for all 16 latents. Conditional on naming, swapping a true gloss for a false one changes it by 0.021 (95% CI -0.11 to +0.15). Truth is not inert: with no token named, the true gloss scores 0.281 higher. A judge-free baseline that marks tokens whose surface form appears in the description reproduces the effect: on the true gloss it matches the judge when the token is named (0.824 against 0.803) and scores zero when it is not (-0.004 against 0.629), and across the 100 latents of our selection pool it tracks the judge at r = 0.61. The judge is therefore not insensitive to semantics; it anchors on the token when one is available. Our false glosses are domain swaps that leave predictive content largely intact, so a high score licenses "this description names the right token" rather than "this description is correct", at least for latents where the token nearly is the explanation.