TARPLE: Target-Aware Semantic Prompt Learning for Few-Shot Object Detection in Remote Sensing
Abstract
Vision-language models (VLMs) align visual and textual representations, yet prior few-shot remote sensing detection methods typically rely on either visual or text features alone rather than jointly leveraging both. We introduce TARget-aware Prompt LEarning (TARPLE), which extends prompt learning to few-shot object detection by learning semantic prototypes adapted to target objects using the vision and text encoders of a pretrained VLM. We further study richer semantic supervision through object descriptions and find that, while class labels alone perform best, description-based prototypes become increasingly effective with more visual examples. Second, we examine robustness to image degradations common in satellite imagery and find that semantic prototypes remain substantially more stable than visual prototypes under corruption. This robustness is associated with improved vision-language alignment, measured by the shift between learned and hand-crafted prompt embeddings relative to target-object images. Across four remote sensing benchmarks, TARPLE consistently improves few-shot detection, with particularly strong gains for fine-grained categories that are difficult to distinguish visually.