PM1: A Multimodal Foundation Model for Genomes, Phenotypes, and Images at Biobank Scale
Abstract
Precision medicine relies on integrating diverse patient data, including genetic variants, clinical phenotypes, and medical imaging, to inform personalized prevention, diagnosis, and treatment. However, learning general patient-level representations from these heterogeneous modalities remains challenging due to their high dimensionality and pervasive missingness. We propose PM1, a multimodal foundation model that learns meaningful patient-level representations across genetic, clinical, and imaging data from 337,129 individuals in the UK Biobank, with a focus on retinal imaging and ophthalmic traits. To address modality dimensionality, PM1 adopts an intermediate fusion architecture with modality-specific encoders and a shared Transformer backbone, where a novel pathway-based genotype encoder enables the incorporation of over 600,000 genetic variants by reducing dimensionality while preserving trait-relevant variation. In addition, a contrastive loss invariant to modality missingness allows PM1 to learn robust patient-level representations despite only 6\% of UK Biobank participants having complete modality coverage. By partitioning modalities into interpretable tokens, PM1 enables interpretability methods to investigate known and potentially novel associations across genetic, clinical, and imaging features, and their interactions. Evaluated across downstream tasks, including phenotype prediction, conditioned image generation, and attribution analyses, PM1 demonstrates the value of multimodal patient representations, improving predictive performance over unimodal baselines and competing methods while enabling interpretable analyses in settings with heterogeneous and incomplete patient data.