You Can't Simulate Grandma: Elder User Simulation Against a Record That Stops at 65
Abstract
Simulated users are validated by comparison to real ones, which presumes a human record at the resolution the simulation claims. We report a case where that presumption fails silently. Six LLMs rate 200 politeness items under age personas from 25 to 85, plus a no-persona control: 28,800 ratings in all. Inside old age they disagree in direction. Between persona 65 and 85, paired across items and Bonferroni-corrected, one model shifts detectably toward more polite and two toward less, so at most half can be describing the same population. We then verify from released raw data how seven datasets record age (six annotation corpora and the PRISM alignment corpus), finding two failure modes past 65. Some stop resolving: POPQUORN compresses everything past 65 into one topless bin, its largest; PRISM stops at "65+"; D3CODE's last boundary is 50; DICES releases only generation labels. Others resolve but run dry: of the three that record exact ages, one holds twelve people aged 75 or over, one holds a single person, and one holds none. No public subjective-judgment corpus we audited distinguishes a 75-year-old rater from an 85-year-old one. Finally we run the control this implies: prompted at eight distinct ages inside the bin, three of six models return identical ratings for a 66- and an 88-year-old on the median item, and none matches the disagreement among five real elders we recruited through community networks. We close with a reporting item for simulation work and a recruitment path that reaches the population platforms do not.