Sparsity, Not Alignment: How a Standard Interpretability Test Manufactures Discovery in a Genomic Foundation Model
Youssef N Zerta
Abstract
Sparse autoencoders (SAEs) are increasingly used to decompose scientific foundation models into features that are then read as mechanisms. We built a causally grounded test of that reading for a genomic foundation model, using saturation-mutagenesis MPRA - which measures the effect of every single-nucleotide substitution in a regulatory element - and aligning each feature's in-silico-mutagenesis sensitivity with the measured per-position effects against a random-direction null with false-discovery control. Applied directly, the test reports a discovery: $18.2\%$ of features clear the null's $95$th percentile against a $5\%$ chance rate, $291$ survive Benjamini--Hochberg correction, and the effect persists under label shuffling, repeat masking, a log-likelihood baseline, element subsampling, three seeds, two SAE families and two architecturally distinct models. \textbf{The discovery is an artifact.} The procedure thresholds one tail but never checks that the excess is one-sided, and it is not: $18.6\%$ of features fall beyond the opposite threshold, the identical FDR procedure applied to the lower tail returns \emph{more} discoveries ($324$ vs $291$), the two distributions differ in scale rather than location ($P(\text{real}>\text{null}){=}0.51$), and the reported effect grows monotonically with dictionary size. What the threshold selects is features whose alignment score is more \emph{variable} than a dense random direction's - which sparser units are mechanically. We give four cheap diagnostics that expose this, plus a noise ceiling for reproducibility claims, and argue the failure mode is generic to interpretability pipelines that threshold a noisy per-feature statistic.
Chat is not available.
Successful Page Load