Beyond Perturbation Robustness: Biological Equivalence in Genomic Foundation Models
Padmaja Mohanty
Abstract
Genomic foundation models are commonly evaluated through predictive benchmarks or perturbation sensitivity, but neither directly tests whether learned representations respect known task-relative biological equivalences. We study translation-preserving equivalence in protein-coding DNA, where synonymous codon substitutions change nucleotide sequence while preserving the encoded amino-acid sequence. Using 91 human coding sequences, we generate 10 independent synonymous transformation paths per source at depths $k=1,2,3$, paired with edit-matched nonsynonymous controls. We evaluate frozen Omni-DNA-20M, Omni-DNA-116M, NucEL-93M, and JEPA-DNA-NTv3-100M using representation geometry, matched equivalence retrieval, and downstream amino-acid-composition stability. Omni-DNA remains at chance on matched retrieval (0.496 and 0.500) despite smaller absolute representation distances at the larger scale. NucEL shows clearer equivalence structure, with retrieval of 0.596 (95% CI [0.572, 0.621]) and a positive downstream stability margin of 0.00382 (95% CI [0.00290, 0.00486]); JEPA-DNA shows weaker evidence. These results suggest that smaller perturbation distances should not be interpreted as evidence of more biologically coherent representations, and that equivalence-based evaluation can complement conventional robustness metrics.
Chat is not available.
Successful Page Load