Correct Scores, Wrong Conclusions: Auditing Forecasting Benchmarks on Chaotic Dynamics
Abstract
Forecast accuracy does not by itself establish scientific forecasting ability. We test five conditions that connect a score to a scientific claim: held-out generalization, complete failure accounting, matched disclosure, structural identifiability, and cross-system transfer. We apply these tests to k-pendulum (k in {1,2,3}) and Lorenz-63 forecasts from language models, a time-series foundation model, learned dynamics, sparse identification, and numerical integrators. Each test changes a substantive conclusion. Learned-model error rises from 0.09-0.63 rad on training trajectories to 0.97-1.14 rad on held-out trajectories, where linearized dynamics and nearest-neighbor lookup perform better. Retaining failed forecasts lowers the reported performance of models with incomplete coverage. An unmatched comparison says that hiding physical constants helps; a paired intervention instead finds a small, model-dependent effect (pooled +0.069 rad; +0.211 to -0.042 across models). Two exact scale symmetries prove that the absolute constants are not identifiable from angle trajectories. Increasing the Neural ODE training set improves error from 1.090 to 0.861 rad but does not reach known-equation integration. The pendulum ordering also fails to transfer to Lorenz-63, where SINDy recovers the polynomial dynamics to numerical precision. On video-tracked hardware the roles invert: the RK4 model with published constants is itself imperfect, and its lead over simple baselines is gone within a second. These results show that a correct leaderboard can still support incorrect scientific conclusions: an unaudited score does not earn trust. We release the prompts, outputs, checkpoints, seeds, and compute records needed to repeat the audit.