Benchmarking Risk Attitudes of LLMs
Bowen Sun ⋅ Rui Min ⋅ Xianyao Li ⋅ Yuxi Wang ⋅ Yang Ye ⋅ Qi Wang ⋅ Eric Du
Abstract
As AI systems are deployed in open-ended, high-stakes domains, a critical behavioral dimension remains unmeasured: whether models differ systematically in how they translate perceived risk into action. We introduce a cross-domain evaluation framework that isolates and quantifies \emph{risk attitude} in large language models (LLMs) by decoupling contextual belief ($B_C$, a model's perceived danger level given situational context) from categorical risk decision ($R_D$, the action chosen in response to that belief). Unlike existing capability benchmarks, which assess factual accuracy but do not probe the belief-to-decision mapping, this framework directly targets the disposition governing whether an agent acts cautiously or aggressively under identical levels of perceived danger. We applied the framework to six frontier LLMs and 100 human participants across spatial navigation, clinical triage, and financial allocation tasks, and identified two qualitatively distinct behavioral profiles. Five models satisfied both reliability criteria (Contextual Belief Consistency and Risk Decision Consistency) across all conditions and tasks. These models showed strong measurement reliability with stable belief formation and belief-to-decision mappings, and exhibited stable, trait-level cross-domain risk attitudes that were preserved across all three tasks (Kendall's $W=1.00$, $p=0.017$). One model, Grok~4, failed both criteria: its belief formation was unstable in Drone Navigation Control task and its decision mapping was inconsistent in Clinical Triage Decision task, yielding task-specific rather than trait-level risk behavior. Across all six LLMs, risk profiles converged toward a restricted region of the broader human risk-attitude distribution, suggesting that alignment training constrains models toward a narrow behavioral consensus. These results demonstrate that the framework reliably measures risk attitude as a stable, model-specific behavioral dimension and can serve as a diagnostic tool for pre-deployment assessment by distinguishing models with coherent risk personalities from those with domain-contingent risk postures.
Chat is not available.
Successful Page Load