A Self-Improving Agent that Grows its Own DSL: Oracle-Free Operator Induction for Exact Question Answering
Abstract
A self-improving agent has to decide which new skills to keep. In practice it keeps whatever earns a high reward, convinces a judge model, or compresses its library. Each of those can be satisfied by a skill that is wrong, which is fine until the answers have to be exactly right. Over a table of transactions a total is either right or wrong, and a skill kept because it scored well can still be wrong. We ask a harder question: can an agent grow what it can compute with no reward, no judge, and no answer key, and still be exactly right? It can. A single frozen model runs two nested loops over a small, exact domain- specific language (DSL) whose operators compose like code under a sandboxed interpreter. The inner loop adds no operators. It poses questions it has not seen, answers each with one expression, and keeps an answer only when a second expression, written independently from the question, agrees on every record. The outer loop adds capability. When the model keeps reaching for something its operators cannot express, it writes a new operator and keeps it only on evidence that needs no answer key: two independently written implementations that agree on every record, and conservation laws checked against operators already trusted (the parts must add up to the whole). On a fully synthetic transaction corpus, the inner loop roughly triples the answerable-question set from a small start, and the effect holds across 13 starting sets. The outer loop runs unsupervised on three disjoint starting sets and extends every one, inventing operators that lie outside the starting closure and span several capability families; a control that adds no operator answers none of the new ques- tions. The grown DSL is cross-checked against an independent, human-authored gold extractor and is exact on every compute type. None of this holds without the admission rule. It admits none of 138 buggy operator variants under an inde- pendent reference, even when the reference is written by a different model family (Gemma-3-27B), while the teacher-free signals the field relies on (admit-if-it-runs, self-consistency, and compression) admit most of them.