From Survey Histories to Interpretable Digital Twins: Auto-Discovering Task-Specific Personas and Behavioral Mechanisms
Abstract
LLM-based “digital twins” aim to simulate how an individual would behave in new environments or respond to novel questions, given some representation of that individual’s prior responses. A common approach constructs this representation from survey transcripts or summaries and evaluates fidelity by predicting holdout responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we ask whether long survey histories can instead be transformed into compact, task-specific personas that remain predictive for unseen respondents and products within the same task while making simulated decisions more interpretable. We introduce an iterative auto-discovery pipeline that constructs persona representations and learns how persona evidence influences simulated responses. In a pricing task, it identifies two behavioral channels, allocates respondent-specific survey evidence to them, and learns when and in which direction each affects purchase predictions. This yields decision-level interpretability: a human-readable mapping from respondent evidence and question context to predicted responses. We evaluate the procedure on the Product Preferences–Pricing block—40 binary purchase decisions for priced products—of Twin-2K-500 using a double holdout of unseen respondents and products. Learning is restricted to responses from 20 development respondents and 20 calibration products, after which the representation and mechanism are frozen and evaluated on 30 unseen respondents and 20 unseen products. The holdout answers come from two survey waves, and every prediction is scored against each wave separately. For GPT-5.4 Nano, replacing the full survey transcript with the learned channel representation and mechanism improves exact accuracy from 0.573 to 0.623 and from 0.558 to 0.626 in the two waves. Among all evaluated conditions, only the learned representation outperforms the full-transcript baseline in both waves with paired 95% confidence intervals excluding zero. Larger and cross-provider simulators—GPT-5.4, GPT-5 Mini, and Qwen3-8B—given the full transcript do not show this pattern, and none has a paired 95% confidence interval excluding zero against the learned representation in either wave. These results show that the learned representation produces consistent gains against human responses collected in both waves while retaining an explicit mapping from respondent evidence to predicted decisions.