External-Phenotype Validation Reveals Meiotic Recombination in a Primate-Divergence Sequence Model
Abstract
Sequence models can predict evolutionary conservation, but predictive accuracy alone does not show that they have learned a biological process. We ask two questions. Can a model trained only on primate-scale sequence divergence recover the long-term evolutionary footprint of meiotic recombination? And is that footprint concentrated in the hotspots that persist over evolutionary time rather than in the transient hotspots that PRDM9 directs? We tested this directly. We trained a 22M-parameter Basenji2-style convolutional network to predict the nonnegative magnitude of primate divergence from 131,072-bp DNA sequences. The target came from the 233-primate subset of the UCSC 447-way phyloP alignment. Held-out Pearson correlations were 0.591 in validation and 0.582 in test. No windows overlapped across splits. When evaluation was restricted to windows more than 50 kb from all training data, the correlations remained 0.585 and 0.574. The model received no recombination maps, hotspot annotations, or motif supervision. Nevertheless, its predictions were elevated at experimentally mapped DMC1-SSDS meiotic double-strand-break hotspots in both held-out splits (+0.026 and +0.015; both confidence intervals excluded zero). These lifts matched the signal in the observed target (+0.027 and +0.016). In the hotspot-centered profile, the predicted focal excess was +0.015 and the observed excess was +0.015; their difference was +0.0004, while matched background was flat. deCODE pedigree crossover rates provided an independent corroboration. PRDM9-directed hotspots relocate every 0.7-1.3 Myr, much faster than primate-scale divergence accumulates. We therefore expected the present-day canonical PRDM9 motif-positive hotspot class to contribute less to the long-term evolutionary footprint. Observed target divergence was greater in hotspots lacking the canonical PRDM9-A motif than in motif-positive hotspots (interval spanning background). The difference was +0.047 [+0.032, +0.063] and replicated in both held-out splits. An independent composition-matched PRDM9-binding partition also placed the excess in unbound hotspots. Attribution separately recovered PRDM9 as a top-ranked genome-wide sequence feature in all eight TF-MoDISco runs despite the absence of motif supervision. These findings are complementary: the model learned PRDM9-associated sequence information, while the primate-scale footprint was carried by the canonical-motif-negative and PRDM9-unbound classes. Together, these results establish an external-phenotype test that converts attribution from motif description into a falsifiable test of biological knowledge.