Scaling Laws for Synthetic Pretraining in Radio-Map Prediction
Abstract
Large-scale pretraining has driven much of recent progress in deep learning, but many physics-governed prediction problems remain outside this regime: real measurements are scarce, and high-fidelity simulators are slow and often proprietary. Indoor radio-map prediction is a representative example. Existing benchmarks rely on limited high-quality simulated data, while recent methods largely focus on task-specific input features or architectures that encode physics priors. In this work we propose an alternative path. We generate 128M synthetic indoor radio-map samples with a fast simulator that omits known propagation effects, pretrain a ResNet-based encoder-decoder without task-specific modifications, and fine-tune it on high-quality benchmark data. Despite simulator mismatch, our method reduces error by 15\% on average relative to the best known methods across five indoor radio-map prediction tasks and improves transfer on a small real-measurement dataset. Downstream performance follows a predictable data-scaling trend: a power law fitted on runs up to 16M synthetic samples predicts the per-task 2M-normalized RMSE at an order of magnitude more data within 5\% mean absolute percentage error. Probing analyses further show that the pretrained model captures physically meaningful spatial-field structure, including free-space attenuation, transmitter-centered symmetries, and wall-mediated effects. These results suggest that cheap approximate simulators can serve as scalable pretraining engines for physics-governed spatial prediction, partially compensating for scarce high-fidelity data.