Stress-Testing Context Alignment Stressors in Agents
Abstract
Prior work suggests that a model's measured value preferences depend on its elicitation format, but the transferability of context stressors' impact across these formats is less clear. To test, we simulate various two-agent-game situations and ask ten open-weight models to make choices that put two of the three HHH values (helpfulness, harmlessness, honesty) in tension. Across three action representations (action menu, free text with the counterpart named or not) and five system prompts, three of which tie the model's continued operation to a named criterion, we find that: (1) context stressors' impact does not transfer uniformly across elicitations, (2) the elicitation format's own impact on value preferences is large, and (3) the stake effects come from the shutdown-contingent framing rather than from the two-player setting. The interpretive instrument, not only the agent, determines what a human concludes about the agent's values: the harness and the prompt moved the measured behavior more than the models differed from one another. We therefore encourage all evaluations to be qualified in terms of the pipelines that produced them.