UncertainGen: Scalable Uncertainty-Aware Representation Learning of DNA Sequences
Abstract
Learning effective representations of DNA sequences is fundamental to genomic data analysis, especially in the presence of high sequence similarity and inter-species DNA sharing. Existing approaches rely on deterministic embeddings, such as k-mer statistics or representations from large language models, which fail to capture the inherent uncertainty in short and ambiguous DNA fragments. We propose UncertainGen, a lightweight probabilistic representation learning framework that embeds DNA sequences as distributions in a latent space. The learned embedding variance captures positional ambiguity, increasing for sequences compatible with multiple genomic clusters. We establish theoretical guarantees on embedding distinguishability, and show that the probabilistic formulation induces a data-adaptive metric that expands the effective representational space. We evaluate UncertainGen in the metagenomic binning task to demonstrate the practical benefits of uncertainty-aware representations. Experiments show that probabilistic embeddings consistently outperform deterministic k-mer and large genome foundation models, while remaining lightweight and efficient.