Solver-as-Teacher: Solver-Guided On-Policy Post-Training Framework for PDE Foundation Models
Abstract
PDE foundation models (PDE-FMs) promise general-purpose surrogate simulation, yet deploying a pretrained model to a new physical regime inevitably introduces distribution mismatch, demanding task-specific post-training. The dominant approach, supervised fine-tuning (SFT), is fundamentally off-policy: the model trains on pre-generated states but is evaluated autoregressively on its own. In chaotic regimes, this mismatch is catastrophic as per-step errors compound exponentially along the Lyapunov spectrum, and exact pointwise trajectory matching beyond the predictability horizon becomes an ill-posed training signal. To resolve this, we introduce Solver-as-Teacher (SaT), an on-policy post-training framework for PDE-FMs. During training, the student rolls out autoregressively; a numerical solver branches multi-step correction targets from each student-visited state; and the student is updated against these solver corrections, learning to correct its own rollout dynamics. We instantiate SaT in a unified post-training suite with SFT, physics-informed temporal alignment, and learned-teacher on-policy distillation. Across seven 2D/3D PDE benchmarks and six pretrained PDE-FMs, SaT variants achieve the lowest 50-step rollout RMSE in 23 of 24 in-distribution model--task pairs. The best SaT variant reduces RMSE over SFT up to 90.0\%, and SaT gives the lowest RMSE on all seven OOD parameter-shift benchmarks.