Long-Context Generation Is a Sampling Problem
Allen Roush ⋅ Minh N Nguyen ⋅ Ravid Shwartz-Ziv ⋅ Judah Goldfeder ⋅ Sanjay Basu
Abstract
Long-form generation in language models often collapses into repetitive loops well before the nominal context limit. This is commonly attributed to model capacity or training data, and addressed through costly retraining or post-training pipelines. We show that the sampling algorithm is a primary determinant of long-context output quality, improving outputs without further training costs. Across ten open-weight models from 1B to 120B parameters, output lengths from 8K to 64K tokens, and more than 13,500 generations, full-distribution-aware samplers (P-less, Top-H, top-$n\sigma$) dominate automated metrics and human evaluation, while standard truncation (top-$p$, top-$k$) collapses into late-context repetition; min-$p$ falls in a middle tier. A 1,500-annotation Amazon Mechanical Turk study finds P-less winning more than nine in ten head-to-head pairings against a top-$p$ 0.9 baseline on a 120B model, with 10-gram repetition approaching zero at 64K tokens. A per-step selected-token rank diagnostic isolates the mechanism: standard truncation drives the rank to the argmax floor by the final quartile, while full-distribution-aware methods stay in a healthy band across the full window. This gap that widens with model scale, establishing sampler choice as an under-appreciated bottleneck for long-context generation in open-weight models. We frame long-context generation as sampling from outside the training distribution, and encourage further study for usecases like long-context agentic coding.
Chat is not available.
Successful Page Load