Stable Encodings, Changing Downstream Sensitivity: Measuring SAE Feature Identity Across Fine-Tuning
Abstract
SAE features are often compared across model checkpoints using activation or decoder similarity, but these measurements do not establish whether a feature retains the same intervention role after fine-tuning. We test this implication by freezing checkpoint-zero SAE coordinates and separately measuring semantic selectivity, activation retention, local downstream sensitivity, and exact intervention effects. In a counterbalanced Gemma experiment, fixed coordinates retain topic selectivity and activation structure while their sensitivity and intervention effects change across learned output mappings. Sensitivity changes far more than activation and accounts for more endpoint-effect variation. An independently trained Pythia SAE reproduces the dissociation. Exact state--parameter hybrids find larger residual-state and interaction components than the parameter-only component in a targeted Gemma replay. Scale controls show that native feature magnitude contributes materially: Gemma's normalized predictive fit fails, while a common-dose Pythia replay retains 18.4\% of the native mapping range. Two emotion panels fail checkpoint-zero qualification, so we do not claim non-topic generality. Stable encoding is therefore insufficient evidence of stable task-conditioned intervention continuity.