A Mechanistic Interpretability Study of an Astronomical Foundation Model
Abstract
Scientific foundation models trained on multimodal data are increasingly deployed across the sciences, but whether their internal computations correspond to known scientific operations or proceed through different mechanisms remains unclear. Understanding these mechanisms is essential for assessing reliability, identifying failure modes, and judging when predictions can be trusted in scientific workflows. We present the first mechanistic interpretability study of an astronomy foundation model, focused on AION-1's multimodal advantage in redshift prediction from broadband photometry and imaging over photometry alone. We find that this advantage concentrates on extended and dust-attenuated galaxies, with the gain localized by layerwise probing to AION-1's first encoder block. We examine two attention heads in this block and find that representation and function dissociate in opposite directions across them. The first head linearly encodes classical photometric features and organizes its image output along a radial gradient, yet ablation and activation patching show this structure has no measurable causal effect on the prediction. The second head shows no distinctive probe signature on the photometric-redshift estimation operations we tested, yet ablating it on dust-attenuated galaxies eliminates the multimodal advantage and amplifies the attenuation-bias signature beyond the photometry-only baseline, indicating that this head implements a correction for attenuation bias. Whether this reflects a methodological limitation in interpretability or an undiscovered astronomical empirical relation is itself an open question, and one that bears directly on using interpretability as a tool for scientific discovery.