Neural RNA splicing models rely on reading frame information unavailable to the splicing machinery
Abstract
Deep neural networks can accurately predict RNA splice sites from genomic sequence, but whether they do so by simulating the natural splicing process is unclear. We evaluate this question by examining mutations that should not have a strong effect on splicing, but are correlated with it in the human genome. We focus specifically on mutations that affect frame alignment because this is known to be relevant to machinery that operates spatially-and-temporally separately from the splicing machinery. We find that predictions made by commonly-used neural tools are impacted far more by deletions or mutations that alter translational reading frame, producing in-frame stop codons, than by those that preserve reading frame. This dependence on reading frame causes these models to systematically miss many poison exons, a clinically relevant category of exons where reading frame is disrupted. This dependence is much weaker in "grey" box models that build in splicing-specific structure. Our results imply that widely-used black box models of splicing rely on information that is not accessible to the RNA splicing machinery.