No Model Required: Text Entropy Rate Filtering Prevents Iterative Fine-Tuning Collapse
Lewis Mitchell
Abstract
Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $H_K$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a superior training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the established gold standard) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $H_K$-filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all $p < 0.001$). We validate $H_K$ as a cross-domain entropy proxy ($\beta = 0.924$, $R^2 = 0.746$) and collapse detector ($\rho = +0.454$, $p < 0.0001$) across 4 domains, 3 temperatures, 2 generator--scorer model pairs, and 1,680 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
Chat is not available.
Successful Page Load