Truth as a Compression Artifact in Language Model Training
Abstract
When transformers train on contradictory data, the same problem with both correct and incorrect solutions, which answer do they prefer? We hypothesize that next-token prediction, as a compression process, favors whichever answer cluster has lower description length; truth benefits only when errors lack internal structure. We test this by training transformers (3.5M-1B parameters) from scratch on controlled corpora, systematically varying the structure of errors. We find that (a) when errors are random, models develop a correctness preference scaling from 65% to 85% with model size; (b) when errors follow a single coherent alternative rule, this preference vanishes (~45-51%); (c) two competing wrong rules suffice to restore it (47% to 78%). The pattern reproduces on Wikipedia paragraphs with entity substitution (71% vs 46%) and at 1B scale on a mixed natural-text corpus (77% vs 47%). These results are consistent with the hypothesis that, in controlled contradictory corpora, model preference tracks the relative compressibility of competing answer systems rather than truth per se.