CultureRed: Benchmarking Culture-Specific AI Safety Based on Global Statutes
Abstract
Safety alignment and evaluation of LLMs has become increasingly difficult due to the diverse definitions of harmful behavior across cultures. Existing safety benchmarks are largely culture-agnostic or built on ad-hoc taxonomies, and mainly evaluate refusals for prohibited requests while overlooking legally defined exceptions, making over-refusal difficult to quantify. To address these gaps, we introduce a large-scale culture-aware data synthesis pipeline and CultureRed as the first global culture-specific AI safety benchmark based on 252 legal statutes across 15 countries. Based on legal statutes, our pipeline automatically transforms statutes into actionable policy rules that cover both prohibitions and legally defined exceptions, cluster policy rules into risk topics and fine-grained categories, and generates policy rule-conditioned data that models should either refuse or comply with. The resulting benchmark contains 54,470 policy rule--prompt pairs spanning 6,256 fine-grained risk categories. We further run a human study with government experts to validate the dataset quality. Leveraging CultureRed, we conduct a comprehensive evaluation for 14 advanced generative LLMs and 16 frontier guardrail LLMs, unlocking a range of findings, such as (1) Culture-specific safety performance varies substantially, with EU presenting the most challenging for generative models (average Harm Rate of 39.4%); (2) Generative model series evolution exhibits divergent patterns, where GPT series achieves balanced alignment (5 times Harm Rate reduction while maintaining Help Rate above 90%), but Claude series shows safety-helpfulness tradeoff (5 times Harm Rate reduction but Help Rate drops from 61.9% to 49.9%); (3) Guardrail models systematically under-guard unsafe content, with high false positive rate.