Guarding the Life Code: Preserving Membership Privacy in Genomic Foundation Models
Abstract
Genomic Foundation Models (GFMs) have become a promising paradigm for decoding DNA sequences, yet their privacy risks are critically underexplored.Trained on human cohorts, GFMs can unintentionally memorize sensitive genomic signatures, that enable membership inference attacks (MIAs) and expose individuals to long-lasting privacy harms.However, existing defenses often fall short in genomics: globally applied mechanisms can substantially degrade motif-sensitive utility, and model-agnostic post-processing provides limited insight into which genomic regions drive leakage. Through a systematic privacy audit, we find that membership leakage in GFMs is highly non-uniform, concentrating on a small subset of high-risk samples and sparse token spans corresponding to biologically meaningful patterns. Motivated by this finding, we propose GenoGuard, a model- and sample-adaptive defense that shifts from blanket protection to targeted, region-aware mitigation. GenoGuard (i) localizes privacy-leaking subsequences via gradient-based token attribution and performs token-level risk reduction through selective gradient routing, (ii) applies sample-level risk reduction with risk-adaptive label smoothing guided by a reference-calibrated loss gap score, and (iii) provides region-level attribution maps for privacy auditing and biological interpretation. Experiments across representative GFMs (Mistral-DNA, Nucleotide Transformer) demonstrate that GenoGuard consistently improves privacy against multiple MIA families while preserving strong fine-tuning performance and interpretability.