Turnip: Tokenizers Are Secretly Context Compressors for Transformers
Abstract
Transformer language models process text as sequences of tokens. Their full self-attention cost scales quadratically with sequence length, making tokenization critical to efficient long-context modeling. Conventional tokenizers such as byte-pair encoding (BPE) require exact reconstruction of the input, which restricts context compression. In this paper, we formulate tokenization as a predictive compression problem and show that lossless tokenization is not necessary for attaining the best achievable next-byte prediction loss. Based on this framework, we introduce Turnip (the neural tokenizer pipeline), a context-dependent neural tokenizer that integrates with existing language models and provides direct control over context compression. On the evaluated models and benchmarks, Turnip nearly doubles the compression ratio of BPE while matching its generation quality.