Auditing Learned Dynamics in Weather Models
Abstract
Learned weather models are experimentally accessible computational systems whose internal representations can be inspected and manipulated. Using sparse autoencoder features from GraphCast, we test whether causal relevance requires physical semantics and whether physically associated latent directions can be steered to improve extreme-event forecasts. First, ablating a feature locked to GraphCast's computational mesh increases global geopotential height error despite lacking conventional physical semantics. Second, suppressing convection- and low-level-vorticity-associated features weakens predicted tropical-cyclone deepening, while amplifying them reduces Hurricane Ida's intensity error by up to 69\%. Together, these results separate physical interpretability from causal utility and establish sparse features as experimental coordinates for auditing and intervening on learned weather dynamics.