Stepwise Benchmarking on Real-World, Data-Scarce Groundwater Problems
Abstract
In scientific machine learning, we use benchmarks to compare models across a variety of physical systems, but we lack methods to reliably assess their scalability to real-world applications. We present a stepwise benchmark that keeps the underlying physics fixed while increasing complexity and spatial extent under realistic, increasing constraints of data scarcity. We apply this concept to 2D subsurface heat transport with interacting open-loop geothermal heat pumps, where long-range advection, heterogeneous subsurface properties, and large spatial domains are characteristic of real applications. The benchmark consists of three steps of training data, ranging from isolated heat pumps up to thousands of interacting pumps in deployment scenarios. The final step is additionally evaluated on an unseen metropolitan-scale domain to assess out-of-distribution scalability. We provide a common evaluation protocol and leaderboard, problem-specific metrics, training and test data, and baselines. Several established architectures fail to predict the source-induced temperature perturbations already in the early steps. ConvNeXt performs reasonably well on the first two steps, while training a problem-specific LGCNN demonstrates that the later benchmark steps are generally feasible. The results show that computational scalability becomes a major constraint when applying scientific machine learning models to large spatial domains. The benchmark complements existing model-comparison suites by addressing the gap between advances in generalizable ML methods for PDE-based applications and their application to real-world challenges.