RadarMAE: Injecting Physical Inductive Biases into Masked Autoencoders for Advanced Radar Object Detection
Abstract
Masked Autoencoders (MAE) have emerged as a powerful self-supervised paradigm in vision, yet their direct application to radar perception remains bottlenecked by a fundamental domain gap. Specifically, while conventional MAEs are optimized for RGB images using isotropic grid patching and standard mean squared error (MSE) loss, radar range-azimuth (RA) heatmaps exhibit sinc-oriented anisotropic spatial structures as well as complex physical artifacts, such as heavy-tailed multipath interference and sidelobe spikes. In this paper, we propose RadarMAE, the first radar-centric MAE tailored specifically to the physical properties of radar signals. We present an axial tokenization strategy that preserves the resolution disparities and continuous sinc-shaped lobes of RA maps. Furthermore, considering that the range and azimuth dimensions exhibit distinct physical artifacts, we decouple their reconstruction objectives by introducing a multipath distribution loss along the range axis to capture heavy-tailed multipath reflections as well as a sidelobe distribution loss along the azimuth axis to learn deterministic sidelobe spikes. Our extensive experiments demonstrate that injecting these domain-specific physical priors enables robust representation learning for radar object detection. Notably, our method achieves a state-of-the-art 64.17% bird's-eye-view (BEV) AP on the K-Radar dataset using only a single-frame RA map setting, significantly outperforming the previous baselines.