Structuring Survey Histories for LLM Digital Twins: A Controlled Study of Persona Representation
Abstract
LLM-based digital twins aim to predict how individuals would respond in new settings from representations of their prior responses. Existing evidence that substantial compression of survey histories need not reduce predictive accuracy suggests that representation structure, rather than information volume alone, may be an important design choice. We study a fixed Background–Decision procedure–Evaluation (BDE) representation that organizes respondent information into identity and knowledge, reasoning procedure, and preferences before simulation. We evaluate BDE on Twin-2K-500 (Toubia et al., 2025), which contains more than 500 survey responses and 88 held-out evaluation items across 17 tasks for 2,058 respondents. In a paired evaluation with 100 respondents, a fully inferred BDE profile improves overall accuracy over the raw survey transcript by 1.27 percentage points (pp) on gpt-5.4-nano (p < 0.01) and 1.19 pp on gpt-5.4-mini (p < 0.001). Three analyses then examine section labeling, extraction depth, and component composition. Deleting all section headers while leaving the profile text unchanged costs 0.52 pp of overall accuracy on nano and gains 0.31 pp on mini, with no statistically significant differences on either. Allowing deeper inference during extraction likewise yields no statistically significant gain on either simulator. By contrast, adding Evaluation to a Background–Decision procedure profile improves overall accuracy by 1.44 pp on nano (p < 0.001) and 1.22 pp on mini (p < 0.001). These results attribute the gain primarily to profile composition rather than section labeling or deeper extraction.