Beyond 3 Million Tokens: A Multi-Modal Foundation Model for Full-Resolution Heliophysics
Sujit Roy ⋅ Johannes Schmude ⋅ Ata A Asanjan ⋅ Thorsten Kurth ⋅ Rohit Lal ⋅ Kshitiz Mandal ⋅ Vishal Gaur ⋅ Harris Abdul Majid ⋅ Nikolaos Dionelis ⋅ Berkay Aydin ⋅ Himanshu Patil ⋅ Andres Munoz-Jaramillo ⋅ Campbell Watson ⋅ Juan Moreno ⋅ Manil Maskey ⋅ Rahul Ramachandran
Abstract
Context length remains a fundamental bottleneck for vision foundation models operating on high-resolution imagery. Current architectures rarely exceed 1M tokens, forcing practitioners to downsample inputs at the cost of fine-grained spatial information. This limitation is particularly acute in heliophysics, where satellite instruments record full-disk solar observations at native $4096 \times 4096$ resolution across $13$ channels, and downsampling discards the small-scale magnetic structures and localized dynamics critical to understanding solar phenomena. In this work, we present the first vision foundation model for heliophysics trained at native 4K resolution on approximately $15$ years of multimodal solar data ($257$ TB) spanning eight extreme ultraviolet channels and five magnetic field and velocity products from the Solar Dynamics Observatory (SDO). We adapt MultiMAE with a dual-view formulation: 25% of tokens are observed directly, while the remaining 75% are replaced with fixed Gaussian random Fourier projections that act as structured, non-invertible frequency views of the masked content. Using two-way parallelism (feature + sequence) with FP8 mixed precision, we scale the context window to $>3$ million tokens at $8 \times 8$ patch size, a $3\times$ advance over prior work. We demonstrate strong scaling efficiency up to $1024$ GPUs across five hardware configurations, with an architecture capable of processing sequences exceeding $10$M tokens. The learned representations achieve strong zero-shot reconstruction across masking ratios up to 90% and full missing-modality scenarios. On downstream tasks (AR segmentation and EVE irradiance prediction) with a frozen encoder, we outperform SOTA by 9.5% and 35.3% respectively.
Chat is not available.
Successful Page Load