Measurable Does Not Imply Transportable: A Prospective Audit of Model Explanations in Multi-Year Field Trials
Abstract
Model explanations are increasingly used as scientific evidence, yet it remains unclear what they license scientists to conclude about a relationship beyond the data on which they were constructed. When a fine-grained explanation fails to replicate, a common diagnosis is that its target was measured too noisily, suggesting that reliable measurement should yield a reproducible explanation. We test this implication directly. We formalize prospective explanation transportability: the extent to which the structure of an explanation learned from past data is recovered in a genuinely future context. This property is distinct from both feature importance and measurement reliability. Using multi-year Genomes-to-Fields maize trials, in which each new year is a real future context rather than a random resample, we evaluate 16 agronomically interpretable environmental concepts across five strictly forward splits and one prospectively locked test year. Coarse environment-level explanations transport reliably (median cosine 0.83); their locked 2024 value, 0.811, fell within a pre-specified 80% prediction interval. For grain yield, genotype-conditioned explanations have a historical median cosine of 0.30 and range from 0.006 to 0.822 across years. Reliability does not order these outcomes: the year with the highest split-half reliability had the lowest explanation transport. To separate prospective transport from estimation variability, we also independently resampled each G×E explanation within a fixed context. Across the 2020–2023 transitions, within-context reproducibility was high (0.77–0.95), whereas prospective cross-year alignment was only 0.01–0.36. A second phenotype transports strongly at both resolutions, and non-temporal validation inflates apparent interaction-level transport by +0.22. These results show that target reliability and within-context explanation reproducibility do not by themselves establish when a model explanation can support a reproducible scientific claim.