Interpretable latents do not imply causal biosecurity use
Abstract
Sparse autoencoders (SAEs) make dense model states inspectable, encouraging the hope that biologically interpretable features will also be causal and useful. We audit this chain in METAGENE-1 using controls for near-duplicate sequences, nucleotide composition, feature search, interventions, operating-point calibration, and taxonomic identity. Two selected viral-function coordinates retain bounded cross-family associations, but a detector restricted to four named coordinates falls well behind raw-state and sequence baselines, especially at a low false-positive operating point. Under taxonomic shift, sequence models retain some global ranking ability but remain near chance at the low-false-positive endpoint. Distinguishing these outcomes shows that a biologically interpretable feature need not provide causal leverage over the parent model or a useful screening signal.