AQBENCH: Benchmarking Neural Surrogates for Air Quality Forecasting
Siddharthan Dileep ⋅ Sanchit Bedi ⋅ Pareshbhai D Parmar ⋅ Ayush Maheshwari ⋅ Sri H Kota ⋅ N M Anoop Krishnan
Abstract
Neural surrogate benchmarks for spatiotemporal PDEs are dominated by idealized test beds with periodic boundaries and synthetic dynamics. Architecture rankings established on these benchmarks do not transfer to real atmospheric chemistry, and standard regression metrics conceal the failure modes that determine operational utility. AQBench evaluates ten architectures across five model families on high-resolution WRF-Chem simulations of $PM_2.5$, $NO_2$, and $CO$ over the Indian subcontinent, against five operationally-grounded objectives: long-horizon stability, exceedance detection, extreme-episode bias, advection-diffusion residuals, and multi-pollutant forecasting. Three findings emerge. Spectral neural operators, the strongest family on canonical PDE benchmarks, rank in the bottom four on 168 hour RMSE here. Models that lead on aggregate regression underestimate during extreme pollution episodes and miss the majority of true exceedance events. Joint training on $PM_2.5$, $NO_2$, and $CO$ improves $PM_2.5$ accuracy across most architectures with no parameter increase, with Transolver gaining 21\% at 168th hour. The dataset, evaluation framework, and trained baselines are released.
Chat is not available.
Successful Page Load