Reliability-Aware Foundation-Model Supervision for Label-Efficient Robotic Perception
Abstract
Foundation models are pretrained predominantly on RGB imagery available online, and rarely see the sensing conditions autonomous robots and systems encounter in deployment, such as thermal imaging. We investigate whether off-the-shelf foundation models, such as SAM3, Grounded-SAM2, and YOLO-World, can act as effective automatic annotators for task-specific object detectors, reducing the annotation burden. We use these foundation models as pseudo-label teachers within the Plug-and-Play Active Learning (PPAL) framework, which selects the most informative samples for annotation across a series of active-learning rounds. To determine when their pseudo-labels can be trusted, we introduce an agreement-based reliability mechanism that combines teacher confidence with spatial agreement between the foundation-model teacher and the task-specific student's predictions. The results suggest that this agreement strategy is an effective criterion for identifying trustworthy predictions in RGB images. We additionally evaluate the framework under an RGB-to-thermal maritime domain shift, where foundation-model pseudo-label quality is substantially lower than on RGB imagery, showing that predictions cannot be assumed to generalize reliably across sensing domains. Overall, reliability-aware integration of foundation-model pseudo-labels, guided by teacher--student agreement, shows promise as a route to label-efficient adaptation of autonomous robots and perception systems.