In-Context Learning for Remote Sensing Vision: A Semantic-Aware Rotation-Robust Diffusion Framework
Abstract
Remote sensing imagery supports diverse visual tasks, such as object detection and semantic segmentation, but these tasks vary across scenes, imaging modalities, and target objects. Since many remote sensing scenarios lack sufficient annotations, conventional supervised learning is often difficult to apply, motivating in-context learning as a flexible alternative.We present an in-context learning framework for remote sensing visual tasks, where a new task is specified by a single annotated example and applied to unseen images without retraining at inference time. Built upon diffusion models, our framework establishes target-reference associations through cross-attention and transfers the structural relationship between the reference image and its annotation to produce task-specific predictions. Unlike natural images, remote sensing objects are often small and provide weak semantic cues in early diffusion steps, making semantic alignment between target and reference images difficult. We address this limitation with a semantic-aware cross-attention that combines semantic and fine-grained details for more accurate cross-image matching. Remote sensing objects also appear at arbitrary orientations, causing large appearance variations and unstable in-context transfer. To improve robustness, we introduce a rotation-robust learning strategy that reduces sensitivity to orientation changes. Experiments show that our method achieves strong one-shot performance, establishing a flexible paradigm for adapting remote sensing models to diverse visual tasks.