Integrating digital twins with randomized experiments
Abstract
Randomized controlled trials (RCTs) identify treatment effects for an enrolled trial population by assigning treatment independently of potential outcomes, conditional on baseline covariates. These trial-specific effects may not generalize to a broader target population when the target and trial covariate distributions differ. This problem becomes harder when individual-level target population data are unavailable. Pre-trained large language models (LLMs) offer one way to address this data limitation by generating digital twins (DTs) with synthetic covariates intended to resemble the target population. Synthetic covariates alone, however, are insufficient, because neither the RCT covariate distribution nor the DT covariate distribution necessarily matches the target covariate distribution. We therefore propose a statistical inference procedure that integrates RCTs with calibrated DTs using external target population summaries. The procedure represents the target covariate distribution as a calibrated mixture of the RCT and DT covariate distributions, allowing DTs to contribute covariate information while limiting the influence of poorly calibrated synthetic data. The procedure also incorporates LLM-generated auxiliary outcome predictions through calibrated outcome regressions, improving precision without changing the estimand. Theoretical results and simulation studies show that the proposed estimators reduce bias relative to RCT-only estimation and improve efficiency for overall and subgroup target treatment effects.