Policy Parity: Cross-Lingual Inconsistency in Policy-Based AI Safety Moderation
Abstract
Safety guarantees are only meaningful for global-scale AI systems if they hold consistently across languages. We study policy parity: whether semantically aligned prompts receive the same moderation decision when the language changes. We evaluate gpt-oss-safeguard-20b under a fixed safety policy on English, Arabic, Spanish, and Hindi prompts. Controlling for stochasticity, we find cross-language disagreement rates of 4.1–5.8%: 9.5–14.9% of prompts judged unsafe in English are judged safe in another language, versus 1.5% in the reverse direction. We then investigate how these disparities propagate when guard-model decisions are used as training labels: language-specific supervision induces significantly larger English-relative false-negative-rate gaps and cross-lingual score disparities than using the guard model's decision on English, especially on unsafe content. These results suggest that preserving semantic consistency in weak supervision can matter more than generating labels in the target language itself.