Cross-seed reproducibility depends on the measurement, not only the model: sparse autoencoder features in a protein language model
Theodore He
Abstract
Sparse autoencoder features are increasingly treated as units of analysis in protein language model interpretability, but whether they survive a change of random seed has not been measured. We train forty sparse autoencoders on one layer of ESM-2 650M, identical except for the random seed. Roughly half of features find a cross-seed counterpart at a cosine similarity of $0.70$, the threshold in common use, though reconstruction quality is nearly identical across seeds. Features aligned with curated biological annotation are neither more nor less reproducible than unaligned ones: the association is bounded at $0.36$ percentage points per standard deviation. The reproducibility figure depends heavily on choices that are rarely reported: it ranges from $80.6\%$ to $8.9\%$ across matching thresholds, and adopting the criterion used by another published study lowers it by nineteen percentage points -- almost the entire difference between our figure and theirs. Reported reproducibility rates are therefore not comparable without the criteria that produced them.
Chat is not available.
Successful Page Load