DiffRisk: Diffusion Representation Learning with Informative Missingness for Health Risk Prediction
Shailesh Dahal ⋅ Ratri Mukherjee ⋅ Nicholas Mathews ⋅ Kishlay Jha
Abstract
Health risk prediction aims to predict a patient's potential health risks (e.g., mortality) using the longitudinal information present in their electronic health record (EHR). While deep learning-based risk prediction models have shown great promise, they are susceptible to the issue of missing data prevalent in real-world EHR. Most prior works mitigate this challenge by treating missing data as an uninformative perturbation to be ignored or imputed. However, inaccurate imputation may lead to the generation of sub-optimal patient representations that adversely impact risk prediction performance. Moreover, recent studies show that missing data in EHR does not necessarily imply error or noisy observation and instead may communicate informative clinical signals useful for health risk prediction. To address this, we propose a novel diffusion-based representation learning approach (namely $\textbf{DiffRisk}$) that explicitly captures $\textit{informative missingness}$ by modeling conditional dependencies induced by partial observations in EHR. Specifically, DiffRisk learns latent representations with coherent conditional structure, where any subset of observed features defines a distribution over the unobserved components. To achieve this, we introduce a conditional score-based objective and approximate it via denoising diffusion that enables efficient learning without explicit density estimation. Empirical results on five risk prediction tasks show that the proposed approach consistently outperforms state-of-the-art baselines.
Chat is not available.
Successful Page Load