Attachability Is Not Portability: A Pretrained Nucleotide n-Gram Memory Does Not Transfer to Frozen DNA Language Models
Abstract
A hashed n-gram memory trained inside a nucleotide language model is trivially easy to detach and bolt onto a different, frozen DNA language model: addressed by hashes of raw nucleotides, it is indifferent to how the host tokenises DNA, and a zero-initialised projection makes the attachment an exact identity at initialisation. We separate two properties this convenience invites one to conflate. Attachability — mechanical reuse across incompatible tokenisers — holds, and cheaply: 2–8% FLOP overhead, no host weight modified. Portability — a benefit attributable to the pretrained contents, holding across hosts — does not. A natural grafted-versus-ungrafted comparison appears to establish portability, but it introduces 1.5–1.8M trainable parameters, so it measures the augmentation package rather than the learned contents. Against a capacity-matched control differing from the treatment in exactly one tensor — the frozen memory table — the apparent host-general benefit disappears: the content effect is heterogeneous in sign across three hosts, and negative on all three for the representation probe the experiment was built to run. The gain also decomposes differently on each host, so pooled comparisons obscure its source, and an apparent depth trend belongs to one host. Across three hosts spanning single-nucleotide, 6-mer and BPE tokenisation, four tasks, three depths and 1,188 fine-tuning runs, the conclusion is not that memory augmentation fails but that neither attachability nor whole-graft improvement establishes portability of learned contents.