Generation in the Limit as a Criterion of Scientific Success
Abstract
Asking whether a generative model is successfully learning rather than reproducing its training data usually means asking whether its outputs are valid and novel. Kleinberg and Mullainathan's generation in the limit for language learning is the idealized form of that standard: a learner succeeds if it eventually produces only valid, previously unseen members of an unknown target language — achievable, remarkably, where identifying the target language is not. We ask what it amounts to as a criterion of scientific success. Recast within the topological framework of formal learning theory, generation in the limit is a formal dual of identification in the limit, a well-studied criterion of scientific success. We show that these two criteria are independent. Additionally, novelty offers two interpretations: the generated evidence may be unobserved or unentailed by what has been observed. We argue that the latter is the relevant notion for scientific inquiry and show that, unlike the former, it is insensitive to how the evidence is labeled and available only for questions no evidence ever conclusively settles.