Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet
Zeyuan Hu ⋅ Noah Brenowitz ⋅ Akshay Subramaniam ⋅ Jaideep Pathak ⋅ Tao Ge ⋅ Mohammad S Abbas ⋅ Suman Ravuri ⋅ Karthik Kashinath ⋅ Noel Keen ⋅ Naser Mahfouz ⋅ Peter Caldwell ⋅ Mike Pritchard
Abstract
Kilometer-scale convection shapes precipitation extremes, tropical organization, and cloud feedbacks, but most global atmospheric models approximate these processes at 25--100\,km resolution. Global storm-resolving physics models resolve convective systems explicitly, but at a cost---roughly one MWh per simulated day on exascale supercomputers---that limits their use for long-duration atmospheric simulation. We introduce STRATA (Storm-resolving Tile-based autoRegressive Atmosphere Transformer Architecture), the first autoregressive AI emulator for global storm-resolving atmospheric dynamics. STRATA is trained on the highest-resolution atmospheric dataset yet used for global AI emulation: 17 days of output from the SCREAM physics model at 4.9-km resolution (${\sim}25$ million grid cells) sampled every 10 minutes. Our central premise is that since on 10-minute timescales atmospheric dynamics are predominantly local, training on small spatial tiles trades scarce global temporal samples for abundant local spatial samples and enables global rollout via overlapping-tile blending. STRATA combines 3D patch embedding and local 3D neighborhood attention for tractable modeling of high-resolution atmospheric tiles, a novel Stereographic Rotary Position Embedding (StereoRoPE) for grid-invariant positional encoding, and a pixel-space de-aliasing decoder that suppresses patch-scale rollout artifacts. An iso-FLOP scaling study reveals that km-scale emulation requires ${\sim}10\times$ more FLOPs per horizontal grid point than coarse-resolution AI weather models, consistent with the higher information density of convective-scale dynamics. Despite training on only 17 days of SCREAM output, STRATA produces stable 24-hour global rollouts with realistic km-scale dynamics across diverse weather regimes, though large-scale biases develop with lead time. STRATA achieves 48 simulation days per megawatt-hour---about 50 times better energy efficiency than the underlying SCREAM physics model---and 741 simulated days per wall-clock day at 512 H100 GPUs. Code and dataset will be publicly released.
Chat is not available.
Successful Page Load