Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks
Abstract
Integer sequences from the On-Line Encyclopedia of Integer Sequences (OEIS) are increasingly used to benchmark mathematical reasoning in language models. What those benchmarks measure is unclear: accuracy on a public, indexed corpus can reflect retrieval rather than induction. We introduce an exactly computable reference learner—two-part minimum description length (MDL) over holonomic recurrences, evaluated on every prefix as terms arrive—and use it to audit the benchmark. Three findings follow. First, MDL difficulty is a parameter count: the discovery point is predicted by a combinatorial identifiability bound on the selected operator, and is invariant to term magnitude. Second, a wilderness regime occurs at OEIS scale: across three runs at a 30-term budget, 47–51% of prefix-fitting sequences have no full-length operator; an all-terms recheck recovered an operator for only 29/1887 examined cases (1.5%). The revision rate is 11–13% under the original verifier on two samples and 29% when singular leading coefficients are admitted. Third, our initial hypothesis that models would confabulate in this regime was not supported: they hedge there, while confident errors concentrate on the easy stratum for all three models. Recognition-conditioned accuracy suggests that accuracy on OEIS benchmarks can reflect recall, and MDL supplies a structural difficulty signal unaffected by whether the sequence appeared in training data, within one hypothesis family. Code, data and all model responses are released at https://github.com/sabilashang/where-induction-runs-out.