Don't Deploy Fine-Tuned Genomic Foundation Models Without Privacy Evaluation: Reconstruction Vulnerability Is Unpredictable Without Empirical Measurement
Abstract
This position paper argues that large language models are widely deployed on sensitive genomic data, both as fine-tuned classifiers for downstream tasks and as an Embeddings-as-a-Service (EaaS) genomic foundation models shared between institutions, without evaluating whether their embeddings can be reverse-engineered to reconstruct the original nucleotide sequences. This absence of evaluation is not benign: for aggregated embeddings, reconstruction vulnerability is architecture-dependent, cannot be predicted without empirical measurement, and is actively increased by fine-tuning for a significant subset of widely deployed architectures. Evidence from independent empirical studies shows that fine-tuned embeddings reduce reconstruction vulnerability for some architectures while significantly increasing it for others, suggesting that models are deployed with privacy properties ranging from improved to degraded, depending on unassessed architectural choices. Per-token embeddings, increasingly shared in EaaS settings, are universally reconstructible across the genomic foundation models evaluated to date; the effect of fine-tuning on per-token vulnerability remains unevaluated. The privacy outcome depends on the interactions among four factors: embedding type, fine-tuning procedure, pretraining and tokenization strategy, as well as sequence length. None of these factors is currently evaluated in practice. We propose a minimum reconstruction-based privacy evaluation standard, organized as three sequential deployment decisions and two attacker tiers matched to research and clinical stakes, and call on the community to adopt it as a prerequisite for deploying any large language model on sensitive genomic data.