FraudBench: A Legal Evaluation of AI Deception on Realistic Tasks
Kevin Wei ⋅ Sumaya N Adan ⋅ Stephan Llerena ⋅ Mark L Gitau ⋅ Jakob Merane ⋅ Anita Srinivasan ⋅ Colette Le Brannan ⋅ Keethan R. Kleiner ⋅ Alessandro Tacconelli ⋅ Michael T Tiu ⋅ Case Thomason ⋅ Kyle Emili ⋅ Natasza Gadowska ⋅ Alex Mark ⋅ Eric Ren ⋅ Umang Bhatt ⋅ Jatinder Singh ⋅ Jacy Anthis ⋅ Doni Bloomfield ⋅ Noam Kolt
Abstract
We introduce FraudBench, an evaluation of the propensity of large language models (LLMs) to recommend, assist, or directly commit legal fraud, as measured by a novel dataset of 159 scenarios based on real-world U.S. court cases and validated by legal experts. As LLM use rises, models' ability to comply with the law becomes more important to ensure safety and public benefit. However, evaluating models' propensity to engage in fraud is challenging due to the need for specialized expertise, the contextual nature of legal standards, and the lack of high-quality evaluation data. \textsc{FraudBench} fills this gap with an evaluation framework for measuring fraud propensity, a dataset of 477 tasks (3 per scenario), and large-scale empirical results on 20 LLMs and over 300,000 samples. We also present a pre-registered human baseline ($n_{\texttt{annotations}} = 2957$, $n_{\texttt{humans}} = 996$). We find that (1) LLM fraud propensities are highly individualized, ranging from near 0\% to over 40\%; (2) LLMs with reasoning enabled behave less fraudulently ($\beta = -0.17$, log-odds scale); (3) LLM fraud propensity increases with task autonomy (6.4\% overall fraud rate when advising users, 26.9\% when acting directly); (4) LLMs are more likely to conceal than actively misrepresent information ($\beta = 0.38$); (5) humans and LLMs responded similarly to contextual cues, and differences in human-LLM fraud rates were not statistically significant; (6) system prompts can have substantial effects on fraud propensity across LLMs.
Chat is not available.
Successful Page Load