Dead Salmons Have a False-Positive Rate: A Calibrated Null-Model Test for Interpretability
Shivam Dubey ⋅ Manan Wadhwa
Abstract
Interpretability methods are used to claim what a model represents. But a method can report a confident finding on a network that has learned nothing. In the neuroscience version of this failure, a dead salmon in an fMRI scanner showed "brain activity" because the analysis lacked a multiple-comparisons control. We ask whether interpretability has the same problem, and measure it. Across six models ($124\mathrm{M}$–$9\mathrm{B}$, five families) and three method families (linear probing, concept-direction analysis, and sparse-autoencoder (SAE) feature selection, including two released production SAEs), a naive significance test flags up to $100\%$ of findings on untrained models: a wrong-baseline failure for a single probe or concept-direction (random features carry lexical structure), a genuine multiple-comparisons one for SAE feature selection. The effect survives regularization and larger samples and is invisible to the shuffled-label control task of Hewitt \& Liang (which passes at ${\sim}2.5$–$7\%$). A calibrated null-model test, comparing the statistic against its distribution on $K$ matched untrained models ($K=1$–$2$ is useless), restores near-nominal false-positive rates, with near-null power retained (measured for the probe). Because the plug-in Gaussian threshold under-corrects on heavy-tailed pools, we recommend a distribution-free empirical-quantile test with $K\geq 20$, whose leave-one-out false-positive rate is ${\approx}1/(K{+}1)$ regardless of pool shape, and validate its control with bootstrap confidence intervals for probes and concept-directions across five families up to $7\mathrm{B}$. The calibration is revealing, not merely permissive: applied to the released-SAE max-over-features statistic it exposes that statistic as carrying essentially no signal under a fair control (the trained model does not separate from random, $z=0.4$). We also catch a dead salmon in our own pipeline: a scoring convention fabricated an apparent "capacity law" that vanished under a fair control. The test is a drop-in control we argue should be standard practice.
Chat is not available.
Successful Page Load