When Does Interpretability Support Discovery? Identification-Aware Audits of Model Internals
Abstract
Interpretability can suggest discoveries about what a model has learned, but a compelling probe, visualization, or activation patch is not yet reliable knowledge. We introduce causal capability tomography (CCT), an identification-aware audit protocol for determining which discovery claim an internal analysis can support. The protocol separates behavioral capability, representational decodability, exact matched transfer, latent-specific transfer, complete-mechanism claims, and external validation. It requires a declared task latent, matched counterfactuals, a behavioral gate, an internal intervention family, same-procedure controls, uncertainty, and explicit failure semantics. A hash-locked ten-seed study on arithmetic carry and two-hop reachability exposes a false-discovery pattern: shortcut-trained models achieve IID accuracy 0.989/0.982 and probe accuracy 0.997/0.871, yet paired counterfactual fidelity is 0.010/0.009 and declared patch accuracy is 0.016/0.018. Each of four fixed paired sign tests gives p = 0.001953. A separately frozen GPT-2-small IOI audit uses 128 discovery pairs and 512 disjoint confirmation pairs. At the discovery-selected layer, bounded patch recovery is 0.780 (95% stratified bootstrap [0.770, 0.791]), above all 99 broad permuted-donor nulls (maximum 0.379; plus-one rank p = 0.010 under the declared within-stratum donor-exchangeability null) and 20 same-norm direction controls (mean 0.047). Because the broad donors do not preserve every lexical nuisance, this establishes exact matched transfer rather than latent-specific attribution or a unique circuit. CCT turns that distinction into a falsifiable standard for deciding when model internals can, and cannot, support discovery.