Sequence-to-Expression Models Fail to Generalise to New Genomic Neighbourhoods
Abstract
Machine learning for synthetic biology is moving beyond interpreting DNA towards designing it. This shift requires verifiers, i.e. models that predict how designed DNA sequences will be expressed before they are constructed, but it remains unclear whether current sequence-to-expression (S2E) models can perform this role outside native genomic contexts. Baker’s yeast provides a uniquely powerful eukaryotic testbed for this question: its genetic tractability enables complex genetic perturbations such as genome-wide reporter integrations and large structural rearrangements that remain difficult to construct at comparable scale in mammalian genomes. Here, we benchmark two S2E models zero-shot across eleven yeast tasks. Both models approximately recover the experimental ordering of short, gene-proximal variants, but perform at chance level on genome-wide reporter integrations and poorly on structural rearrangements. For native-promoter swaps, their predicted transcript-abundance ranges are only approximately two-fold, whereas measured fluorescence spans one to two orders of magnitude more for two reporter genes. These results identify contextual generalisation and dynamic-range compression as major obstacles to using current S2E models as verifiers for genome design. An open-source implementation of our benchmark can be found online: https://github.com/anon-schmoo/neurips-icbinb-2026