When Consensus Is Not Correctness: Auditing Cross-Seed Agreement in Learned Chaotic Simulators
Abstract
Does agreement across random seeds improve the reliability of learned chaotic simulators? We audit CoRE, a convex readout method that penalizes predictive disagreement among four frozen random reservoirs. Across five synthetic, fully observed deterministic systems and 1,536 held-out forecasts, matched controls separate the effects of consensus from joint fitting, feature capacity, recurrent coupling, and scalar shrinkage. At the adequate working points of Lorenz-63 and forced Duffing, positive consensus shows no supported improvement in forecasting or long-run distributional fidelity over Joint-0, the matched fit without the consensus penalty. On Lorenz-96, a recurrently block-diagonal control with matched total state count outperforms both the independent ensemble and the dense large reservoir, showing that the gain does not require cross-block recurrent coupling. In the under-tuned Kuramoto-Sivashinsky (KS) case, CoRE, stability-selected ridge, and norm-matched ridge each keep 5 of 8 rollouts below 50 training standard deviations, compared with 0 of 8 for Joint-0. All four methods yield 0 of 8 below 20 standard deviations, and every KS forecast fails before one Lyapunov time. The apparent stability gain therefore reflects amplitude control under a permissive threshold, rather than a consensus-specific improvement in dynamical fidelity. These exploratory results identify confounds in attributing reliability gains to consensus; they do not establish that consensus fails for well-tuned simulators in general.