Searching Beyond Human Anomalies: Expected Utility Violations in Large Language Models
Johnathan Sun
Abstract
Behavioral evaluations of large language models (LLMs) often reuse experiments designed to expose departures from rational choice in humans. We show that this can miss where a model's own violations occur. LLM choices are correlated with human choices across a large dataset of risky choices, but modally reproduce none of 60 expected utility anomalies discovered and validated using human data. We therefore train neural networks to predict LLM choices and use them to search for problems where LLMs exhibit violations of expected utility. Searches yield problems based on common-ratio and monotonicity effects that induce generalizable deviations from expected utility far larger than human-derived anomalies.
Chat is not available.
Successful Page Load