A Mechanistic Interpretability Study of an Astronomical Foundation Model
Abstract
Scientific foundation models trained on multimodal data are increasingly deployed across the sciences, but whether their internal computations correspond to known scientific operations or proceed through different mechanisms remains unclear. Understanding these mechanisms is essential for assessing reliability, identifying failure modes, and judging when predictions can be trusted in scientific workflows. We present the first mechanistic interpretability study of an astronomy foundation model, focused on AION-1's multimodal advantage in redshift prediction from broadband photometry and imaging over photometry alone. We find that this advantage concentrates on spatially extended and dust-attenuated galaxies, with the multimodal gain localized by layerwise probing and block-resolved patching to AION-1's first encoder block. Within that block, how well a property is linearly decodable does not predict how far it can be localized. Spatial extent is causally used, and one head dominates linear probing for it, yet ablating that head has negligible effect on the prediction: the representation is not localizable to a single attention head or linear direction. Dust attenuation is only weakly decodable anywhere in the encoder, yet one head is necessary for the dust-specific multimodal advantage, and ablating it leaves the model more attenuation-biased than the photometry-only baseline. What that head computes remains unidentified: none of the astrophysically motivated operations we tested accounts for it. Whether this reflects a methodological limitation in interpretability or an undiscovered astronomical empirical relation is itself an open question, and one that bears directly on using interpretability as a tool for scientific discovery.