How LLMs Distinguish Threats from Offers
Abstract
Threats and offers are fundamental to strategic interactions between agents. Competently navigating multi-agent settings requires the ability to distinguish between the two, reason about the intentions behind them, and understand how other agents assess them. With the prospect of LLM-based autonomous agents being deployed as representatives or advisors to humans in consequential domains, it is therefore important to understand how LLMs reason about threats and offers. In this paper, we propose a formal theory of threats and offers grounded in causal games, and use it to empirically investigate how current LLMs classify and reason about such proposals. We find that LLMs' threat/offer classifications align with our formal definition, with each model's threat labels tightly tracking its own judgement of whether the proposal leaves the recipient worse off. LLMs also predict their co-players' classifications accurately, though in our setting this accuracy is not clearly above what is already achieved by simply projecting their own labels onto the co-player.