HELICS: Biobank-scale Conditional Synthetic Genome Generation via Latent Flow Matching
Ahmad Abdel-Azim ⋅ Xihong Lin
Abstract
Individual-level genotype data underpin genetic discovery, disease-risk modeling, and therapeutic development, yet many critical settings remain data-limited, including underrepresented ancestries, rare or imbalanced disease cohorts, and restricted-access biobanks. Conditional synthetic genome generation would support privacy-aware benchmarking, simulation, and data augmentation, but requires learning an ultra-high-dimensional discrete distribution over hundreds of thousands to tens of millions of correlated single-nucleotide polymorphisms (SNPs), or sites of DNA variation. We introduce $\mathbf{HELICS}$: $\mathbf{H}$igh-fidelity g$\mathbf{E}$neration via $\mathbf{L}$atent $\mathbf{I}$nterpolant $\mathbf{C}$onditional $\mathbf{S}$ynthesis, a phasing-free latent flow matching framework for genome-wide SNP-level generation from biobank-scale genotype data. HELICS tokenizes chromosomes with convolutional autoencoders, reducing $>$$600{\rm K}$ SNPs to $4{,}096$ continuous latent tokens, and trains a transformer parameterized conditional flow matching model over the concatenated multi-chromosome latent representation. Trained on roughly $428\rm K$ genomes in the UK Biobank, HELICS generates genome-wide synthetic cohorts without haplotype phasing and achieves stronger fidelity to held-out real genotypes than existing haplotype- and genotype-based simulators, including the highest validation allele-frequency agreement ($R^2=0.992$) and superior preservation of local and long-range SNP correlation structure. Conditioning on ancestry and genetic liability summaries enables targeted generation across genetic risk profiles; in perturbation analyses, by increasing the conditioned disease trait liability, HELICS induces SNP-level dosage changes in generated genomes that are highly correlated with published association weights. HELICS provides a scalable generative modeling framework for SNP-level genome data and a path toward conditional synthetic cohorts for data-limited genetic studies.
Chat is not available.
Successful Page Load