SARLAND: When Spatial Detail Matters in Very-High-Resolution SAR
Abstract
Public very-high-resolution (VHR) synthetic-aperture radar (SAR) data are becoming easier to access, but it remains unclear when finer spatial information improves downstream prediction – and whether “finer” is a property of the image grid at all. We introduce SARLAND, a new multi-modal dataset built from SARLO-80 and S1GRDLC resources: 67,083 standardized 1024 × 1024 Umbra crops from 1,406 acquisition passes, paired with Sentinel-1 and WorldCover and extended with Overture building footprints, supporting land-cover, building-footprint, occupancy and scalar targets on one cohort. A four-epoch screen over 35 model configurations gives practical baselines. We then compare two ways of removing spatial information at the same nominal scale. Resampling rendered Umbra amplitude to a 10 m grid changes ResNet101 land-cover mIoU from 0.4526 to 0.4508, and a Fourier control on the same rendered image gives 0.4513 – both within seed-level variability. Restricting spatial-frequency support in the complex SLC before amplitude rendering, at a nominal ≈ 9.9 m response width, reduces three-seed mIoU from 0.4547 to 0.2545. These operators act at different processing stages despite similar nominal scales, and we argue that the contrast is consistent with differences in how many independent looks survive each. Against our pre-specified expectation, crops dominated by large building footprints – not small ones – show the clearest reference-image advantage (+0.0348 pooled IoU with R50, +0.0255 with R101); a retrospective median-footprint Q4–Q1 contrast remains positive after adjustment for building count and built fraction (+0.0132 [+0.0042, +0.0222]). Nominal spatial scale alone does not specify what has been removed from SAR data, or how strongly a model depends on it.