Causal heteroscedastic structure learning from incomplete temporal data
Abstract
Discovering causal structures in multivariate time series (MTS) is critical in domains such as finance and neuroscience, yet remains difficult in practice due to three intertwined obstacles. First, MTS are routinely corrupted by missing values, and naive impute then discover pipelines can distort the underlying data distribution, leading to false causal graphs. Second, recovering instantaneous links is already hard on complete data, and missingness further entangles imputation with graph learning. Third, real world data often exhibits heteroscedastic noise, whose variance depends on both instantaneous and lagged causes, violating the assumptions of standard causal discovery algorithms. Existing methods fail to address these challenges concurrently, typically assuming complete data, homoscedastic noise, or purely time lagged relationships. To address these gaps, we introduce Cheesefill, an optimal transport framework for causal structure learning under missingness. Cheesefill parameterizes causal mechanisms with conditional normalizing flows to capture heteroscedastic noise, and jointly learns a stochastic correction map that refines naive imputations into trajectories consistent with the inferred mechanisms. Theoretically, we show that optimizing over stochastic correction maps is equivalent to a conditional kernel reformulation of the Kantorovich problem; the neural parameterization used by Cheesefill then yields a tractable objective. Extensive experiments on synthetic and real data show that Cheesefill consistently outperforms existing baselines across different missingness mechanisms.