Learning Shared Latent Cellular Representations for Multi-Assay Prediction
Abstract
In drug discovery, evaluating the biological response of candidate compounds often requires experimental screening of large numbers of compounds across broad panels of biological assays, which can be time-consuming and costly. Consequently, increasing efforts have sought to predict assay responses using scalable cellular phenotypic and transcriptional measurements. However, these multi-modal observations are often incomplete and are not directly aligned across modalities, posing challenges for robust prediction. We propose a multi-modal framework that learns a shared, perturbation-responsive cellular representation and refines it through either a deterministic or diffusion-based transition, complementing existing fusion strategies. Experiments across 156 cell-related assays show that the proposed framework outperforms uni-modal and late-fusion baselines under complete observations, while diffusion-based refinement provides greater robustness as observations become increasingly incomplete. These results support shared representation refinement as a practical strategy for multi-assay prediction from multimodal perturbation profiles.