Culture-Dependent Refusal Disparities in Large Vision-Language Models
Abstract
As large vision-language models (LVLMs) are increasingly deployed in real-world environments, it is critical to ensure that their safety guardrails are robust to the wide variety of visual contexts which they will encounter. An LVLM's propensity to refuse potentially harmful prompts (i.e., its refusal response) is a central component to its overall deployed safety behavior. Despite this importance, the sensitivity of a LVLM's refusal response to visual contexts across different cultures has been relatively understudied. To address this gap, we evaluate the refusal behavior of popular LVLMs under counterfactual changes to cultural contexts depicted in input images. Specifically, we conduct comprehensive experiments investigating how the refusal rate of LVLMs differs for benign and harmful prompts across counterfactual image sets which depict the same person in different religious, national, and socioeconomic contexts. Using a two-stage approach, we first screen the refusal behavior of six LVLMs to identify model-prompt pairs which exhibit a high degree of refusal variability and then conduct a second stage of generation for identified prompts to robustly estimate the magnitude of cultural context refusal disparities. Across a total of 26.3 million analyzed LVLM generations, we identify replicable context-driven effects for both over-refusal of benign prompts and under-refusal of harmful prompts across all six evaluated LVLMs.