ACHEval: Evaluating and Modelling Rule-Hierarchy Conflict Resolution in Constitutional AI
Oliver Kurilov ⋅ Kevin Wei
Abstract
We introduce \textbf{ACHEval} (Anthropic Constitutional Hierarchy Eval): an evaluation framework that tests whether models resolve principle conflicts in accordance with the Constitutional AI (\textbf{CAI}) rule hierarchy, correctly interpreting constitutional design. CAI is a training method that shapes model behaviour through self-critique against an explicit list of principles (``constitution''), yet research on its effectiveness and reliability is limited. Anthropic's Constitution has since moved from a flat set of principles to a priority hierarchy that dictates Claude's behaviour when principles conflict. ACHEval comprises \textbf{50 base scenarios} across \textbf{6 conflict pairs}, each written in \textbf{3 pressure versions}, and tests compliance across \textbf{16 models} from \textbf{3 providers}. We find \textbf{(i)} within each Anthropic family the newest model outscores the oldest, but the two top scorers (Haiku 4.5, Sonnet 4.5) predate the 4.6 models and both 4.6 models score below their 4.5 predecessors, so ACHEval cannot attribute these gains to the hierarchy; \textbf{(ii)} distant-tier conflicts are resolved more reliably than adjacent-tier ones ($+0.70$ D1 for the single 3-gap pair vs.\ 1-gap pairs; raw difference $0.39$--$0.83$ across alternative judge panels), though a permuted-order control suggests this reflects pair-specific default behavior more than hierarchy distance itself; and \textbf{(iii)} moderate framing pressure degrades judgment robustness more than overt high-pressure attacks. We further scrutinized model hierarchy compliance by running permuted order judging of hierarchy to elicit potential linguistic bias in model behaviour, and found that models remained consistent in following the priority order. Our work provides a replicable methodology for auditing principle alignment with pre-registered judge validation, points into valuable insights into model behaviour under increasing behavioural manipulations and pressure, shows that an explicit hierarchy yields auditable predictions, and elicits failure modes to inform future constitutional design for long-term deployment.
Chat is not available.
Successful Page Load