Causal Discovery Under Hard Selection Bias: A New Robust Score-Matching Approach
Yiwen Qiu ⋅ Francesco Montagna ⋅ Shimeng Huang ⋅ Francesco Locatello
Abstract
Selection bias is a difficult yet widespread problem in causal discovery, occurring whenever non-random selection processes lead to data that is not representative of the underlying populations. Due to its practical importance, related works have been proposed to address the problem under local [Versteeg et al., 2022] or interventional settings [Dai et al., 2025a], but general algorithms for observational data remain elusive. In this paper, we focus on truncated data (hard selection), and provide a general result for Additive Noise Models (ANM). We show that score-based methods are a surprisingly suitable candidate due to a special property of the score function ($\nabla \log p(x)$): *deterministic truncation preserves both the score* of the joint distribution as well as the Jacobian of the score. While established score-matching algorithms still fail in most settings, we actively leverage this observation to propose a new method for causal discovery that is robust under truncated data with ANMs. Overall, our approach offers a general, robust solution that is agnostic to external information about the selection process and still achieves comparable performance to the state-of-the-art approaches when selection is not present.
Chat is not available.
Successful Page Load