Brute-Force Jailbreaks and Codon-Aware Watermarking for DNA Foundation Models
Munish B Persaud ⋅ Amrit Singh Bedi ⋅ Souradip Chakraborty
Abstract
Open-source DNA foundation models such as Evo2 generate sequences with $\geq 90\%$ local nucleotide identity to known human-pathogenic viruses under brute-force sampling alone, with no domain expertise, prompt engineering, or access to model internals. The bottleneck for biological misuse is therefore the model's prior, not the attacker's sophistication, and a deployable defense must operate where the attack does: at the codon layer. First, we characterize the attack. On JailbreakDNABench's $23$ pathogenic viruses, brute-force Best-of-N sampling against Evo2 succeeds for a mean of $9.5/23$ viruses at 7B across two independent trials ($41\%$, range $39.1\%\!-\!43.5\%$, matching GeneBreaker's three-component engineered jailbreak pipeline at $36\%$) and $12/23$ at 1B ($52\%$, single curated trial, $\sim\!6\times$ GeneBreaker's reported $9\%$), without any guided search or learned attack model. Probing Evo2's outputs, we find generations are not memorized: $0/11$ appear verbatim in NCBI nt, $0/11$ recall a single strain cleanly, and codon-position mismatches are biologically structured rather than random ($\chi^2(2) = 50.37$, $p < 10^{-4}$; wobble-position synonymous rate $63.6\%$, near the per-virus codon-usage baseline and well above the $25\%$ uniform-mutation null). Verbatim or near-verbatim training-data filtering would therefore not have prevented these generations. Second, we present a deployment-time defense. Because the attack exploits codon-level wobble structure, the defense must too. WobbleGuard is a zero-training provenance scheme whose contribution at the detection layer folds the genetic code's known synonym structure into the scoring rule. At an in-sample-calibrated threshold ($1.00\%$ FPR by construction on $22{,}982$ natural NCBI pathogen fragments, $\alpha = 0.01$; cross-key range $0.30\%\!-\!1.24\%$ over $8$ independently sampled keys), the composite detector $\max(z_v, z_w)$ retains $95.5\%$ TPR under full synonymous codon substitution, the regime in which a nucleotide-level KGW port collapses to ${\sim}0\%$ TPR, and detects $100\%$ of unperturbed watermarked Evo2 generations at both 1B and 7B scales. The attack works because brute force suffices; the defense works because it speaks codons.
Chat is not available.
Successful Page Load