Model Evaluation for Spatially Structured Data: A Case Study of Satellite Altimetry Surface Reconstruction
Xiaoyu Xia ⋅ Abani Patra ⋅ Beata Csatho
Abstract
Machine learning methods used in surface reconstruction from satellite altimetry are commonly evaluated on randomly withheld observations. On an exact-repeat orbit, however, withholding points or one dated pass can leave earlier measurements of the same ground track within meters of the queries. Such a test measures reconstruction with a previously observed track, not interpolation into a gap between tracks. We compare pass and track holdout on ICESat-2 elevation data from Greenland. On Tile19, two operational protocols give inverse-distance-weighting errors of 0.142 and 46.03m, although the protocols induce different query sets. In three regions where the queries are fixed, the track-to-pass error ratio ranges from 13.7$\times$ to 1.0$\times$. After holding out complete tracks, method rankings still change across regions, track assignments, and aggregation rules. Kriging has the lower error on smooth interior tracks, whereas a learned residual model reduces the largest errors in rough margin terrain. Controlled neighborhood experiments further show that context size and spatial coverage account for part of the difference in this setting. For surface reconstruction, evaluation should reflect performance in regions without direct observations, rather than only recovery near previously observed tracks. Our results support holding out complete RGTs and reporting both query-pooled and track-level error separately.
Chat is not available.
Successful Page Load