Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
Hyunsik Kim ⋅ Youngmoon Jung
Abstract
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for large language models (LLMs) in multilingual settings because they cover all Unicode text. Under UTF-8, however, many scripts start from a higher byte-level fallback cost than English: when no learned merge covers a character $c$, the tokenizer must emit multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor: $$ f(c) = |\\mathrm{utf8}(c)| \\in \\{1,2,3,4\\}. $$ A higher floor inflates token budgets, shrinks usable context, and increases per-request cost. Changing the underlying text encoding can reduce this gap, but a single global encoding can also make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer with deterministic per-character routing: $$ r(c) = \\begin{cases} \\mathrm{UTF\\text{-}8}, & |\\mathrm{utf8}(c)| \\leq 2,\\\\ \\mathrm{UTF\\text{-}16}, & |\\mathrm{utf8}(c)| > 2. \\end{cases} $$ Characters satisfying $|\\mathrm{utf8}(c)| \\leq 2$ stay on the UTF-8 path, while characters satisfying $|\\mathrm{utf8}(c)| > 2$ are routed through UTF-16. This reduces token costs for Basic Multilingual Plane (BMP) scripts whose characters typically satisfy $|\\mathrm{utf8}(c)| = 3$ and have the highest token premiums, without raising costs for already-efficient spans. UBE changes only the byte representation presented to byte-pair encoding (BPE): the BPE merge rule remains standard, and exact decoding $\\mathrm{decode}(\\mathrm{encode}(s)) = s$ is preserved. Across intrinsic tokenization evaluations, UBE reduces cross-lingual dispersion in English-normalized token-count ratios $\\pi(\\ell) = m(\\ell)/m(\\mathrm{en})$, lowering cross-script token-budget disparity. In multilingual language model experiments, UBE preserves LM quality within seed variability. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and also slightly lowers English token counts; these reductions translate into more usable context under fixed token budgets and faster content-matched prompt processing.
Chat is not available.
Successful Page Load