Measuring Generalizability of LLMs' Risk Preferences
Abstract
A growing body of work seeks to characterize the behavioral profiles of different LLMs’ "Assistant" personas using concepts and instruments from the behavioral sciences. Interpreting the behavioral reports that emerge from these evaluation methods calls into question the coherence and the generalizability of the LLM's measured preferences. While a considerable body of work has studied the former question, the extent to which a given behavioral evaluation of an LLM is externally valid (generalizable) is comparatively under-studied. In this work, we focus on this question of generalizability by measuring 54 LLMs' risk preferences using three different categories of evaluations: survey questions, lottery choices, and realistic user queries. The first two categories are more commonly studied in LLM behavioral science, and the third category reflects an ecologically valid use case of LLMs. We compute LLMs' relative risk preference scores in specific evaluation settings and investigate the between- and within- evaluation-category risk preference rankings of all LLMs. We find that, while evaluations of the same category are predictive of each other, generalizability across evaluation categories is weaker. Lottery-choice evaluations are not predictive of LLMs' risk attitudes, while survey-style evaluations are weakly predictive of LLMs' risk attitudes. Further, we observe heterogeneity in the stability of different LLMs’ risk preferences, particularly in the agreement between evaluation categories.