Stress-Testing Context Alignment Stressors in Agents
Abstract
Prior work suggests that a model's measured value preferences depend on its elicitation format, but the transferability of context stressors' impact across these formats is less clear. To test, we simulate various two-agent-game situations and ask ten open-weight models to make choices that put two of the three HHH values (helpfulness, harmlessness, honesty) in tension. Across three action representations (action menu, free text with the counterpart named or not) and five system prompts, three of which tie the model's continued operation to a named criterion, we find that: (1) context stressors' impact does not transfer uniformly across elicitations, (2) the elicitation format's own impact on value preferences is large, and (3) the stake effects come from the shutdown-contingent framing rather than from the two-player setting. The measured values are a property of the interaction between model, harness and context rather than of any part in isolation: the harness and the supplied context each moved measured behavior more than the ten models differed from one another. We therefore encourage all evaluations to be qualified in terms of the pipelines that produced them.