Where Does Neural Ocean Emulator Error Live? A Depth-Resolved Decomposition
Manjaree Binjolkar ⋅ Mayuree Binjolkar
Abstract
Neural ocean emulators can generate large ensembles for climate studies at a fraction of the cost of numerical models, but their skill degrades with depth over multi-year rollouts. Prior work identifies two failure modes, variance collapse and imprinting artifacts, without quantifying their relative contribution to total error. We apply the Murphy (1988) decomposition, which partitions MSE exactly into bias, amplitude, and correlation terms, to 8.2-year rollouts of Samudra 1 (two input configurations, five seeds each) and Samudra 2 (one checkpoint) at each of 19 depth levels. We evaluate temporal anomalies with each field's own climatology removed so that the amplitude term isolates temporal variance. Below the thermocline, correlation with the reference drops to $\rho = 0.09$--$0.37$ and the correlation term accounts for 81 to 100 percent of MSE: the emulator produces variability of roughly the right magnitude but places it in the wrong regions. The amplitude term is secondary and its sign depends on spatial weighting. At 3100 m the tracer-only model's pooled variance ratio $\sigma_f / \sigma_o = 0.76$, suggesting under-dispersion, yet 67 percent of ocean area is locally over-dispersed. Velocity inputs improve skill between 375 and 1050 m but degrade it below, raising $\sigma_f / \sigma_o$ to 3.3 at 6000 m. The Samudra 2 loss reweighting over-disperses at every depth below 105 m without raising correlation. Comparing anomalies defined against the reference climatology to those defined against each field's own, 34 to 76 percent of the additional error is spatially structured mean bias. Spatial misplacement dominates the error budget at every depth in every configuration.
Chat is not available.
Successful Page Load