The Tokenization Trap: Domain-Adaptive Byte-Pair Encoding Restores Language-Like Token Statistics for Protein and Genome Language Models
Abstract
Protein and genome language models (pLMs/gLMs) have not exhibited the smooth, predictable scaling behaviour that drives progress in natural-language models. We argue that a central, under-examined cause is tokenization. The near-universal choice of single-residue tokenization (one token per amino acid or nucleotide) collapses the token-frequency distribution onto the marginal residue frequencies, producing a flat, non-Zipfian spectrum (fitted exponent α ≈ 0.4– 0.6). Natural language, by contrast, exhibits a Zipfian distribution (α ≈ 1) whose heavy tail is precisely what neural scaling laws exploit. We formalize this as the tokenization trap and test it with two complementary experiments on real Swiss-Prot sequences: (i) a distributional analysis showing that domain-adaptive byte-pair encoding (BPE) moves the token-frequency exponent into the language- like band (α ≈ 1.0–1.2) with high goodness-of-fit, while single-residue tokeniza- tion and an English-trained GPT-2 BPE do not; and (ii) a controlled language- modelling experiment in which, at a fixed parameter budget, domain BPE re- duces bits-per-residue by∼9% relative to single-residue tokenization, whereas GPT-2’s English BPE is worst. We further show that GPT-2-on-biology is not even power-law distributed (fit r2 ≤ 0.75), so any single Zipf exponent for it is an artifact, and we adopt a goodness-of-fit-gated robust estimator to report it honestly. A model-size sweep on real Swiss-Prot makes the trap quantitative: single-residue tokenization is nearly flat with scale (bits-per-residue power-law slope b ≈ 0.0025), whereas domain BPE turns parameters into compression about twice as fast (b ≈ 0.0058) and closes the gap monotonically. A frozen-feature linear probe on a composition-controlled protein-family task further shows single- residue features are near chance (0.23 vs. 0.17) while domain BPE recovers the motif structure (0.86). Finally, moving from synthetic probes to a real bench- mark, we couple each tokenizer to an external 21-trait microbial-phenotype task over 364 native genomes: domain BPE restores the Zipfian exponent on real genomic DNA (α : 0.31 → 1.12), wins the matched single-nucleotide contrast at every taxonomic level, and, at∼3 M parameters, matches a 7- billion-parameter single-nucleotide genome model (Evo2), corroborating the trap where it matters most. We conclude that domain-specific subword tokeniza- tion (not merely any subword scheme) is a necessary but not sufficient condition for language-like scaling: it fixes the input distribution, but the transformer itself still only learns correlational patterns. We argue the remaining bottleneck is archi- tectural and outline a spectral, mechanism-aware design (Spectral Composition Attention) that biases the model toward the data’s coupling spectrum rather than surface co-occurrence