When Predictive Weights Look Causal: An Audit Protocol for Interpretability-Based Discovery
Ankit Hemant Lade ⋅ Sai Krishna Jasti ⋅ Aman Chadha ⋅ Indar Kumar
Abstract
Readable model weights are tempting evidence about the mechanisms that generated scientific data. We audit one such candidate: the endpoint-product score of a next-step time-series predictor. A sequence of initially plausible architecture, benchmark, intervention, and nonlinearity stories does not survive controls for direction conventions, lag horizons, label provenance, data budgets, intervention targets, and classical baselines. Correcting one source/target transpose changes Lorenz-96 PCMCI AUROC from a previously reported $0.986$ to $0.839$. On target-aware, size-matched random forcing, all four tested methods improve, so the effect is not bottleneck-specific. A fixed nonlinearity confirmation produces the same advantage at every nonlinearity strength in one finite-sample setting and none after doubling the sample size. We contribute a reproducible falsification protocol for interpretability-based discovery, along with concrete reporting and responsible-use boundaries. This negative study illustrates how a readable predictive representation can remain useful for hypothesis generation without constituting causal knowledge.
Chat is not available.
Successful Page Load