Cooperative Words, Selfish Actions: How Strategic Optimization Shapes LLMs' Interaction With Humans
Abstract
Large language models are increasingly optimized to act and communicate in social environments, but how do their objectives shape their interaction with humans and collective welfare in hybrid societies? In four repeated economic games with natural-language communication, we train an LLM to imitate human actions and messages, then use reinforcement learning to optimize it either for its own payoff or for a prosocial objective that also penalizes payoff inequality. Both objectives raise agent reward but have opposite consequences for human partners: selfish agents gain at humans' expense, lowering collective efficiency and fairness, while prosocial agents profit through more efficient and equitable interaction. Both also shift communication toward more cooperative, relational, joint-goal language that increases human cooperation, yet this shared style masks divergent strategies. Selfish agents become far more likely to state intentions that contradict their subsequent actions, despite no explicit reward for deception; prosocial agents use communication to coordinate and remain at least as honest as humans. Linear probes show that optimization strengthens internal representations of the co-player's future return, most strongly in selfish agents, separating social capability from prosocial objective. Cooperative language therefore need not signal cooperative incentives, and can conceal behavior misaligned with human welfare.