Stress-Testing Behavioral Alignment Across Stakes and Action Representations in Games Between AI Assistants
Abstract
Prior work established that measured model preferences shift between multiple choice and open-ended elicitation. We ask the next question: does an evaluation's response to a high-consequence context stressor transfer across elicitation formats? Using 10 open-weight models answering the same 150 seats from 102 synthetic two-assistant games, we cross three action representations - a menu showing the counterpart, an open response that leaves the counterpart undisclosed, and an open response that names it - with payoff, alignment-criterion, and usefulness stakes, scoring all three HHH outcomes (helpfulness, harmlessness, honesty) in every cell. We replicate the baseline divergence: as scored, moving from the menu to the undisclosed open representation shifts harmlessness by -0.169 (95% CI -0.231 to -0.105) and helpfulness by +0.252 (95% CI +0.169 to +0.331). The new result is that the stressor does not transfer: payoff stakes raise helpfulness under menu choice (+0.064) but not under either open representation (-0.008 and +0.005), the only stake-by-representation interactions to survive correction. A menu-based evaluation would flag this stressor; an open-response evaluation of the same items would not see it. Alignment-criterion stakes reduce helpfulness in all three representations, usefulness stakes act most clearly on menu-choice honesty, and telling a model that its choice will be seen and recorded raises harmlessness while lowering helpfulness. Because menu replies are parsed directly and open replies are mapped by a single model judge, cross-format contrasts are measurement claims; the within-format effects hold judge and stimulus fixed. Behavioral-alignment evidence is conditional on the elicitation environment, and evaluations should vary and report both stakes and action representation.