Evaluation of a Cross-Domain SSL-Pretrained Point-Cloud Encoder on Aerial LiDAR
Abstract
Each point-cloud domain still gets its own model, trained from scratch on its own labels. Point-cloud self-supervision has followed scaling within domains rather than across them. Utonia is a step towards one encoder for all point clouds, pretrained jointly over five domains; remote sensing is one of them, but as aerial photogrammetry: no aerial LiDAR is in its pretraining mixture, and neither benchmark used here appears in it. Is a domain-specific encoder still necessary? Keeping Utonia’s 137 M-parameter Point Transformer v3 frozen and training only a linear probe, decoder probe, or LoRA, we measure how well it segments two large-scale aerial LiDAR (ALS) benchmarks from different countries and sensors, DALES and FRACTAL, and what that saves in labels. On FRACTAL we reach 0.807 mIoU / 0.963 OA on the full official test split using 3.6 % of the training area, exceeding the published fully-supervised RandLA-Net baseline (0.772 mIoU/0.961); on DALES, with the full training area, we match KPConv’s published mIoU (0.815 vs 0.811).