ByteDistill: Cross-Tokenizer Distillation via Chunk-wise Byte-Level Distribution Alignment
Yang Chen ⋅ Xianqi Yu ⋅ SHAOWEI YAO ⋅ Fuyu Lv ⋅ Dan Ou ⋅ Haihong Tang
Abstract
Token-level distillation assumes a shared categorical support, an assumption violated when teacher and student tokenizers segment the same text differently. Existing cross-tokenizer methods sidestep this mismatch by aligning token spaces, matching likelihoods along tokenizer-specific paths, or filtering shared spans, all of which can approximate the original distributional objective or discard probability mass. We argue that decoded UTF-8 bytes are common observable events shared by every text tokenizer, while token boundaries are tokenizer-specific latent transitions. Building on this view, we introduce \emph{ByteDistill}, which compares the conditional distribution of the next observable byte rather than aligning token identities. Its loss, \emph{Byte-Level Distribution Alignment} (BLDA), projects each native softmax into a 256-way observable next-byte distribution inside aligned decoded chunks, marginalizing token-ending mass as a latent boundary transition along the observed tokenization path. ByteDistill consistently outperforms supervised fine-tuning and prior cross-tokenizer baselines in both off-policy and on-policy settings, with off-policy training improving GSM8K accuracy by $1.75$ percentage points over the current state of the art.
Chat is not available.
Successful Page Load