Composition Without Grammar: Next-Token Models Miss Regulatory Structure in Bacterial DNA
Abstract
Natural selection jointly shapes bacterial genes and their cis-regulatory contexts. Context-aware regulatory-sequence design therefore requires more than generating species-typical DNA: upstream regions must encode valid positional grammar and, where relevant, be compatible with the target coding sequence. Because genomic language models are trained on complete genomes, one might expect them to capture these dependencies. We assess whether these models capture local, position-dependent promoter grammar and statistical associations between coding sequence (CDS) and its upstream region. In a comparative analysis of singleton orthologues, upstream similarity increased with CDS relatedness, whereas matched non-homologous controls showed little association. Motivated by this observation, we evaluate three tasks in E. coli and B. subtilis: promoter-element placement, CDS-conditioned upstream generation, and retrieval of the native upstream among 99 decoys. Across five autoregressive models, promoter placement was comparable to shuffled controls. After fine-tuning Evo with a species token, generated sequences reflected the base composition of the requested species and included a Shine–Dalgarno signal. However, replacing, shuffling, or removing the CDS produced only small changes in the generated upstream. Lastly, in the retrieval task, the best autoregressive model ranked the native upstream first in 6.3% of cases. By contrast, the contrastively trained C3P model reached 82.7% on E. coli, which was excluded from its training set. Under the evaluated settings, next-token models therefore capture broad composition more readily than motif identity, spacing, or gene-specific association, whereas a contrastively trained model can exploit the pairwise signal.