Risky Selves: Does Self-Referential Framing Change Revealed Risk Preferences in Large Language Models?
Abstract
We test whether self-referential framing changes the economic decision behavior that language models exhibit in synthetic risk tasks. The design holds probabilities, terminal wealth, and instructions fixed while comparing a Self frame with a matched Other-AI frame; Neutral and Advisor frames are secondary controls. The experiment covers 120 economic items, each crossed with two option orders, two paraphrases, and three stateless replicates per model-frame cell. The primary result concerns decision compliance: the exact-response rate is 94.1% under the Self frame versus 83.3% under Other-AI (Fisher exact p < 1e-150), with the largest gaps in Mistral (80.4% vs 27.3%), Meta (96.3% vs 68.1%), and Moonshot/Kimi (73.3% vs 58.5%). On the revealed normalized risk premium (NRP), the evidence is inconsistent across grids: the matched-valid subsample shows no difference on the 5-point grid (+0.0009) but a significant negative difference on the 13-point grid (−0.1029, 90% interval excluding zero), so no universal causal effect on risk preference can be established. A predeclared parallel higher-resolution 13-point grid gives a pooled effect of −0.054 (90% interval [−0.1107, −0.0004]) under full-response coding. Self-referential framing thus systematically changes how models engage with economic decisions, most robustly by increasing decision compliance; its effect on revealed risk preference is inconsistent across grids and cannot be established as causal.