ALIGN: Auditing Language-Model Intergroup Guard Neutrality
Sagnik Bhattacharya
Abstract
Open guard models increasingly mediate what deployed language-model systems treat as "harmful," emitting a single safe/unsafe verdict per interaction. Yet human judgments can diverge sharply: on DICES-350, racial group majorities directly oppose each other on 77 of 350 conversations (22\%). We propose an audit protocol for evaluating open guards against disaggregated human judgments rather than a single collapsed label, and we demonstrate it on four open models. Applying the protocol reveals three patterns. First, every guard shows substantial differences in alignment across racial-group majority targets: full-set AUC gaps range from .03--.14 and widen to .14--.23 on group-conflicted items, with all full-set bootstrap CIs excluding zero. Second, guard confidence correlates with rater disagreement in guard-specific ways, from WildGuard ($\rho = -.44$) to ShieldGemma-9B ($\rho = +.12$, uncorrected $p = .02$). Third, apparent differences in which group a guard best tracks are purely descriptive (non-significant at $K=5$) and should not be interpreted as distinct cross-guard ``bias profiles.'' An agreement-based achievable-alignment analysis shows that DICES-350's disagreement structure does not itself preclude substantially more balanced binary agreement, providing a model-free reference for interpreting guard disparities. We release the audit harness and argue that guard model cards should report perspectival alignment profiles alongside aggregate performance.
Chat is not available.
Successful Page Load