The Spread and Self-Erasure of LLM Vocabulary in Scientific Writing
Abstract
Quantifying how large language models changed scientific writing requires markers that stay valid as the community learns to avoid them, and measurements that survive the dominant confound in PDF-derived text, which is extraction itself. On a decade of arXiv papers scored under dual extraction, we find that model-associated vocabulary diffuses through the whole writing population, not a machine-written subpopulation. The same authors use substantially more marked vocabulary after 2022 against a negative pre-LLM placebo, and every percentile of the distribution rises. The most recognisable markers meanwhile decay under detection pressure. Famous ``AI tells'' lose most of their rate within two years while frequency-matched unpublicised terms hold steady. A sentence-rhythm detector returns a corpus-scale null dominated by extraction drift, yet on mined disclosure and leakage labels it identifies machine-drafted papers where vocabulary cannot. The two families measure different events. Vocabulary tracks what a population read. Rhythm tracks who drafted the sentences. Conflating them builds detectors whose terms cancel. We release per-paper scores, mined labels, and rewriting ablations.