The Floor Moves With the Probe: Trivial Baselines for Self-Supervised Learning on Satellite Imagery
Abstract
Self-supervised objectives for satellite imagery are compared to one another and almost never to a trivial baseline. We added two: patch-mean raw spectral bands, which has zero parameters, and a random-initialised encoder. The conclusion changed. On PASTIS crop segmentation, under a frozen probe with epochs, per-step and effective batch, and probe budget matched across cells but compute deliberately not matched, the floor is not a constant: the raw-band baseline moves from 4.76 to 21.45 mIoU, a factor of 4.5, as the probe’s receptive field grows from 1 to 9. The margin of the best objective over it moves with it. Temporal latent prediction beats the raw-band floor by 10.56 ± 1.12 mIoU at receptive field 1 and 4.78 ± 1.30 at 3, but by −0.15 ± 1.28 at 5, where the zero-parameter baseline reaches 20.24 mIoU against the encoder’s 20.09, and −1.57 ± 1.28 at 9: its entire advantage over a baseline with no encoder and no parameters is closed by giving the probe a 5×5 convolution. Of the other four, three sit more than two standard deviations below the binding floor and one is indistinguishable from it. We therefore report rankings only at a stated probe capacity, with floors at that same capacity. These are frozen-probe results at one training budget, and we report the compute each cell used.