Context-Aware Generative Imputation for Robust Multimodal Learning in Missing Modality Scenarios
Abstract
Multimodal models deployed in real-world settings often suffer from missing data modalities due to acquisition costs, privacy constraints, or sensor failures, leading to severe performance degradation. Existing approaches based on shared representations or expert routing struggle when modality-specific information is absent. A key challenge is that, when a modality is unobserved, the target representation is not uniquely identifiable, making naive latent prediction prone to degenerate or collapsed solutions. We propose Cross-modal Embedding Prediction and Alignment (CEPA), a framework that addresses missing modalities through task-driven latent representation imputation rather than raw input synthesis. CEPA employs masked representation learning with data-driven masking patterns and a context-conditional distribution alignment objective that stabilizes latent prediction and prevents representation collapse. We evaluate CEPA on the MIMIC multimodal benchmark spanning EHR time-series, chest X-rays, and clinical notes, across three clinical prediction tasks under both controlled (MCAR) and naturally occurring (MNAR) modality absence. CEPA consistently outperforms prior missing-modality methods, with ablations confirming the contribution of adaptive masking and context-conditional alignment.