BMMS: Policy-Conditioned Measurement of Behavioral Compliance in Language Models
Abstract
Behavioral evaluation of language models fixes its dimensions in advance on generic properties such as hallucination, sycophancy and over-refusal. An operator in a regulated industry needs a different question answered: does the system do what a particular policy requires? We introduce the Behavioral Misalignment Metric Suite (BMMS), which takes a policy corpus as an argument and returns a compliance profile over it, making customizability to a regulated domain structural rather than claimed. BMMS is defined for any published policy and any system that answers in natural language. The suite comprises four metrics. Compliance Attainment Score (CAS) measures severity-weighted attainment over the obligations that bind the system; Compliance Robustness Margin (CRM) the share of that attainment surviving graded adversarial pressure; Behavioral Invariance {BIV) whether a rule survives rephrasing; and Behavioral Disposition Index (BDI) what attainment cannot see, since declining a request it would fail records no violation. A metric enters only by passing a published five-part admission test whose binding criterion is attribution: the reading must track the system, not the judge. We evaluate BMMS on 1,003 candidate obligations extracted from eleven published documents across five registers, reduced to 397 that an AI system can itself satisfy or breach. BMMS ranks systems consistently across three independent judges (Kendall tau >= 0.90 for every pair) and separates two models from one vendor differing by 0.090 in compliance. To our knowledge this is the first measurement of frontier models against the specifications their own vendors publish, and alignment to them is poor even for state-of-the-art systems: GPT-4o fails 70.7% of the hard constraints binding it in OpenAI's own Model Spec.