Characterizing Underrepresentation in Generalizing Causal Survival Estimates
Bolun Liu ⋅ Sean McGrath ⋅ Yiren Hou ⋅ Elizabeth Stuart ⋅ Harsh Parikh
Abstract
Randomized trial findings are routinely generalized to broader target populations to inform decisions in medicine and public policy. When the trial sample is misaligned with the target population, generalized effect estimates can yield suboptimal decisions. We identify and characterize the target subgroups that are underrepresented in the trial cohort in time-to-event settings. Here, underrepresentation arises in two distinct ways: baseline underrepresentation, when a subgroup's covariate profile is rare in the trial relative to the target, and time-varying underrepresentation, when differential censoring erodes a subgroup over follow-up. Existing diagnostics either conflate these mechanisms or address only the baseline case. We show that the asymptotic variance of the efficient generalizability estimator admits a decomposition in which only two terms can diverge as representation weakens: a sampling-weight term capturing baseline misalignment and a censoring-weight term capturing differential loss to follow-up. Underrepresented subgroups are therefore precisely those the trial estimates with poor precision, and the decomposition reveals which channel is responsible.We exploit this operationally in $\textbf{TRACE}$ ($\textbf{T}$emporal $\textbf{R}$ashomon $\textbf{A}$nalysis of $\textbf{CE}$nsored data), which combines a self-normalized variance objective with a tree-based search to localize underrepresented subgroups, attribute each to baseline mismatch or censoring, and pinpoint the follow-up intervals where censoring drives the variance inflation; the joint problem is solved by a parametric dynamic program nested inside a Rashomon ensemble of near-optimal trees. On synthetic benchmarks and a case study generalizing the National Job Training Partnership Act (JTPA) Study to U.S. JTPA-eligible adults, the framework recovers underrepresented regions, correctly attributes each mechanism, and identifies the follow-up windows driving temporal underrepresentation; refining the trial and target samples to reliably generalizable subgroups substantially reduces generalized effect-estimate variance.
Chat is not available.
Successful Page Load