In & Out: Social Dilemmas With Frames and Self-Negotiated Contracts
Xuanqiang A Huang ⋅ Charlie Tharas ⋅ Samuele Marro ⋅ Van Q. Truong ⋅ Bernhard Schölkopf ⋅ Emanuele La Malfa ⋅ Zhijing Jin
Abstract
AI agents are becoming increasingly autonomous and, similarly, act in more and more complex environments. Often, these environments are formally underspecified: instructions are imprecise, actions are not known $\textit{a priori}$, system-wide effects are unpredictable, and goals are abstract or underspecified. In this paper, we characterize the cooperation between different agents in such environments within \textit{social dilemma} scenarios and prove an inefficiency theorem of contracts in such complex environments. We further evaluate two orthogonal categories of solutions to the collaboration problem grounded in different economic theories of cooperation, namely contracts and "we"-frames. This frame allows different agents to reason cooperatively and find the best individual mutually beneficial action for the group. We study the effect of a range of contract representations, from natural language to formal contracts and assess the implications on agent cooperation. We observe that contracts of any kind perform 15.33\% worse in environments where agents selfishly choose compared to when they instantiate the group frame; and 41.76\% worse in common-pool resource environments. However, in both cases, having contracts is better than not having any. Our results suggest both self-negotiated contracts and "we"-frames improve cooperation over normal LLM behaviour, with "we"-frames being more robust in complex settings.
Chat is not available.
Successful Page Load