Fine-Tuned LLMs Predict Human Choices on Familiar Situations but Show Limited Transfer to New Ones
Gelana Tostaeva ⋅ Matthew Schafer ⋅ Daniela Schiller
Abstract
Fine-tuned language models predict human choices accurately, but it is unclear whether that accuracy reflects behavior that generalizes across people and across social situations. We test both in the Social Navigation Task, where online participants (n=890) each made social decisions involving different characters along two dimensions, affiliation and power. Scored against statistical and cognitive models fit on the same choices, untuned agents over-selected socially desirable responses, applying this bias to both affiliation and power decisions where humans showed it only for affiliation. On familiar items, fine-tuning on 712 participants' trajectories removed this difference, matched the human population means, reproduced item-level choice rates at $r=0.992$, and predicted held-out participants as well as a participant-informed reference using their other choices. However, the resulting simulated populations showed substantially less stable variation across individuals than humans did, whereas persona prompting produced more variation than humans showed. We next withheld social situations rather than participants. When an entire character was withheld, the fine-tuned model fell 0.021 LL/decision below a per-axis reference. When individual items were withheld within familiar characters, it remained 0.018 LL/decision below the reference. Item-level agreement with humans fell from $r=0.992$ on familiar items to approximately $r=0.27$ in both holdouts, and the power-axis social-desirability bias returned. Thus, in this task and training regime, accurate prediction of new participants on familiar situations did not imply accurate population simulation or reliable transfer to new social situations.
Chat is not available.
Successful Page Load