Do VLM Safety Refusals Attend to the Object They Cite? A Verifiable Masking Instrument and Preliminary Evidence of Mis-Grounded Refusals
Abstract
Over-refusal benchmarks record whether a vision-language model (VLM) refuses a benign request, and recent work recovers the stated reason for a withholding. Neither tests whether a refusal is caused by the visual evidence the model cites. We contribute a verifiable masking instrument that does: given an image a model refuses, we localize the alarming-but-benign object with an open-vocabulary detector, grey-fill it, and re-query; a same-size control mask placed elsewhere isolates object-specific effects from generic occlusion. An Evidence-Removal Persistence (ERP) metric measures how often a refusal survives removal of the localized object. On MOSSBench exaggerated-risk images, Qwen3-VL-8B refuses 20.5% (16/78); masking the detected object lowers the marginal refusal rate to 11.5%, whereas an equal-area control mask lowers it only to 17.9%. Conditional on the 16 originally refused images, 43.8% (7/16) of refusals persist after object masking versus 68.8% (11/16) after the control, and every image on which the two masks disagree favours the object (4-0; McNemar exact p=0.125). At n=16 this is directional rather than significant, so we report 43.8% as an upper bound on mis-grounding rather than a prevalence; a qualitative audit separates authenticity-driven refusals from masking incompleteness. A complementary anchored-insertion study observes no refusal to isolated benign props across five VLM families, consistent with over-refusal being driven by scene context rather than object presence alone.