Composable Causality: A Toolkit for Systematic Time-Series Causal Discovery and Treatment-Effect Benchmarking
Abstract
Real-world multivariate time series datasets rarely come with verified ground-truth causal structure, since the underlying mechanisms are typically unknown and randomized intervention is infeasible in most domains. Empirical research on time-series causal discovery and treatment-effect estimation therefore depends on synthetic data, where the data-generating process is specified by construction. Existing data generation tools fall into two categories. Some are fixed datasets that provide pre-generated time series, which cannot be modified after release. Others are script-based generators, where each script fixes one combination of mechanism, confounding, and missingness, so any new combination requires writing new code. Neither supports the controlled experiments that method development demands, where a single property such as sparsity, noise distribution, or missingness mechanism is varied while others are held fixed. A further gap concerns counterfactual data. Prior work in time-varying treatment-effect estimation has relied on per-paper simulators with hardcoded mechanisms, while general-purpose time-series generators produce observational data only, leaving evaluation of conditional average treatment effect estimators reliant on conditional expectations rather than ground-truth individual effects. We introduce DTCD (Datasets and Toolkit for Causal Discovery), a data generation library that produces multivariate time series together with their ground-truth contemporaneous and lagged adjacency matrices. Topology, mechanism, noise, observational mask, and intervention policy are independent components that compose freely, enabling single-axis sensitivity studies through configuration alone. Paired counterfactual twins are produced by caching the stochastic realization during the factual rollout and re-injecting it during the interventional rollout, recovering individual treatment effects to numerical precision. To demonstrate the library, we present a reference benchmark of eight discovery algorithms across eighty-one discovery configurations and four treatment-effect estimators across eighteen CATE configurations, exposing failure modes not visible in aggregate metrics, including the inability of standard discovery methods to handle missing values regardless of whether the missingness mechanism is informative.