GeoRefusalBench: A Diagnostic Refusal Benchmark for Remote Sensing Vision-Language Models
Yuqiu Li ⋅ Jian Xue ⋅ hao wu ⋅ Ke Lv
Abstract
Remote sensing vision-language models (RSVLMs) are typically evaluated solely on answer correctness. However, Earth-observation images often lack the visual evidence required to answer specific questions. For instance, the ground sampling distance might be too coarse, clouds or shadows may obscure the target, or a single image can't capture temporal changes. In such cases, a reliable model should refuse to answer and provide a valid reason. Failing to do so can lead it to fabricate confident answers, thereby misleading downstream analyses and decisions. To evaluate this refusal capability in the absence of essential geospatial information, we present GeoRefusalBench, comprising 587 episodes drawn from five optical and two SAR source datasets. We construct this benchmark by applying perturbations, which derive from a 83-factor taxonomy spanning five refusal categories and three intensity levels, to solvable binary existence questions. The annotation is hybrid: 193 samples are grounded by pixel and geometric measurements, while 394 rely on consensus among independent verifiers. Evaluating seven open-source RSVLMs under two output protocols yields a peak accuracy of $47.8\%$, which falls below the $50.3\%$ baseline of always answering. A clustered bootstrap analysis confirms this gap is statistically significant. Moreover, high binary refusal $F_1$ scores coincide with predictions collapsing into a single refusal category. The best recall for specific refusal reasons is $28.8\%$, achieved by a model that refuses $94\%$ of all questions.
Chat is not available.
Successful Page Load