Constraint Retrieval Is Not Constraint Enforcement in Large Language Models
Abstract
Prompt-level instructions are the main way users give language models preferences, safety boundaries, and content rules. But a model can remember a constraint and still violate it when generating text. This points to a gap between two abilities that are often treated as one: retrieving a constraint from context, and enforcing it during token selection. Across five model families (0.5B--32B parameters), we find a clear behavioral split: constraints that admit an alternative generation plan are largely followed, while token-level prohibitions fail whenever the prohibited continuation is locally dominant. Mechanistic analyses show that retrieval heads still attend to the constraint at the violation step, but the resulting signal either paradoxically enhances the prohibited token's logit or suppresses it too weakly to change the argmax. Decoding-time interventions can eliminate surface-form violations, but semantic compliance still requires higher-level verification. These results establish that constraint retrieval and constraint enforcement are separable capabilities, and that reliable token-level compliance requires explicit veto mechanisms beyond prompting.