Does Cultural Knowledge Survive Recursive Synthetic Training? Evidence from English and Burmese
Abstract
Recursive training on model-generated data can distort the data distribution and reduce representation of low-probability content, a phenomenon associated with model collapse (Shumailov et al., 2024). Recent work also frames iterative self-training as cultural evolution, motivating questions about which knowledge survives repeated synthetic generation (Guo et al., 2026). We study this issue for culturally specific knowledge in English and Burmese. Using 100 paired Myanmar cultural concepts, we construct a controlled head/medium/tail distribution and recursively regenerate synthetic cultural corpora for three generations with Qwen2.5-1.5B-Instruct. We evaluate both behavioral access through 100 paired multiple-choice questions and corpus-level retention using multilingual semantic similarity. Behavioral accuracy declines from 61% to 54% in English and from 48% to 44% in Burmese under synthetic-only recursion. Corpus retention at similarity ≥0.60 also declines but unexpectedly remains higher in Burmese at G3 (73%) than English (60%). The contrast shows that retention in generated text and downstream access to cultural knowledge need not move together and cautions against treating lexical/semantic corpus survival as equivalent to model knowledge preservation.