A Simple Unigram Cross-Entropy Lens on the Lexical Imprint of Pre-training Data
Jeonghoon Kim ⋅ Woojin Chung ⋅ Woomin Song ⋅ Cheonbok Park ⋅ Nako Sung ⋅ Jinwoo Shin
Abstract
How much of zero-shot benchmark accuracy is captured by word-level alignment between a pre-training corpus and the benchmark target? We study this question through word-level unigram cross-entropy (UCE), a tokenizer-agnostic, model-free measure of lexical alignment between a reference corpus and a target benchmark. Likelihood training allocates probability mass according to the pre-training corpus, and likelihood-based zero-shot evaluation partly reads out the lexical alignment between this mass and the benchmark target. Cleanly tracing how the lexical imprint of pre-training data scales with corpus--benchmark alignment requires controlled from-scratch pre-training that varies corpus identity while holding the training recipe fixed. Across this controlled sweep, spanning $11$ zero-shot benchmarks, $4$ pre-training corpora, and $3$ model scales, we find that corpus--benchmark lexical alignment, UCE, consistently tracks zero-shot accuracy on every benchmark, with corpora more aligned with the benchmark (lower UCE) achieving higher accuracy. The same signal extends from measurement to control at finer scopes. At the document level, strengthening the imprint by selecting training documents to maximize lexical alignment with the benchmark yields a data-selection rule that improves accuracy over same-size random sampling and an importance-resampling baseline across sources and scales. At the prompt level, matching the imprint per prompt by routing each prompt to the drafter whose domain corpus is most lexically aligned with it gives a training-free routing rule for speculative decoding and improves throughput over a generalist-drafter baseline. Taken together, pre-training corpora leave a measurable lexical imprint on zero-shot behavior, and UCE provides a model-free probe of lexical alignment between reference and target data.
Chat is not available.
Successful Page Load