Joint Treatment Effect Estimation from Incomplete Healthcare Data: Temporal Causal Normalizing Flows with LLM-driven Evolutionary MNAR Imputation
Olivia Jullian Parra ⋅ Sara Zoccheddu ⋅ David C Cerezo ⋅ Tom Forzy ⋅ Franziska S Ulrich ⋅ William Sutcliffe ⋅ Jakob M Burgstaller ⋅ Oliver Senn ⋅ Patrick Owen ⋅ Nicola Serra
Abstract
Target trial emulation (TTE) provides a framework for answering causal questions using observational data when randomized controlled trials (RCTs) are infeasible. However, standard methods for treatment effect estimation have been developed in isolation, failing to jointly address the compounding challenges inherent to analyses of observational data such as electronic health records (EHRs). In particular, these challenges include time-varying confounding and missing-not-at-random (MNAR) missingness reaching 50\%-80\% for critical biomarkers. To address this gap, we propose a two-stage pipeline that jointly handles MNAR missingness and causal structure across time. CausalFlow-T, a Directed Acyclic Graph (DAG)-constrained normalizing flow with Long Short-Term Memory (LSTM)-encoded patient history, performs exact invertible counterfactual inference, eliminating the approximation errors and confounding biases where existing variational and adversarial methods fail silently. Ablations on four synthetic and one semi-synthetic dataset with known counterfactuals confirm that its two core design choices address strictly non-overlapping failure modes (DAG constraints for confounding separation, exact inference for structural propagation) with neither compensating for the absence of the other. To handle the incomplete data CausalFlow-T receives as input, we propose an LLM-driven evolutionary imputer and evaluate it with three LLM backends, including two open-source models. Across 30\%-80\% MNAR missingness, the imputer achieves the best pooled rank across biomarker and causal metrics, leading on point-wise accuracy and temporal extrapolation while maintaining average treatment effect (ATE) recovery where statistical methods progressively degrade. Applied to a cohort of adults with type 2 diabetes in Swiss primary care initiating a GLP-1 receptor agonist or SGLT-2 inhibitor, the pipeline recovers a per-protocol weight-loss difference of $-0.98$ kg [$95\%$ CI $-1.01$, $-0.96$] favoring GLP-1 receptor agonists, consistent with RCT evidence and estimated directly from realistically incomplete real-world EHR data.
Chat is not available.
Successful Page Load