AtmoZero: Self-Play Post-Training for Caption-Free Weather Time-Series Captioning
Abstract
Multivariate weather time series are a central modality for forecasting and decision support, and the surrounding language increasingly mediates how numerical weather is consumed by users and downstream models. However, existing pipelines depend on closed-source language model ensembles with expert quality control or on retrieving human-curated text, while prior caption-RL rewards optimize cycle similarity or downstream-task utility, neither of which can tell a captioner that a confidently asserted phenomenon is contradicted by the data, so the absence of an executable, station-observable, and refutation-aware reward signal has blocked the application of self-play RL to this scientific modality. To address this, we propose AtmoZero, a post-training framework that trains a weather-time-series captioner without a single human-written caption, by aligning a language model with a library of seven station-observable scientific verifiers, pairing them with an explicit refutation channel that penalizes caption claims the data contradicts, and combining them with auxiliary reconstruction and forecaster-utility rewards under group-relative policy optimization. Extensive empirical results on a twelve-year ERA5 reanalysis corpus and real-world surface observations show that AtmoZero attains the lowest unsupported-claim rate and forecasting error at every horizon, outperforming closed-source captioners, dense-reward RL, and far-larger foundation models, with extensive analyses confirming that these gains arise from multiple complementary signals and are robust under distribution shift.