LinguaInterpreter: An Auditable Meta-Agent for Autonomous ML Engineering
Abstract
We present LinguaInterpreter, a meta-agent that takes a natural-language machine-learning task and a data directory and produces a submission autonomously, coordinating 77 specialised LLM agents through a hierarchical task tree under a global time budget. Its organising principle is that no claim made inside the system is accepted without the evidence that would make it checkable. Four mechanisms implement this. Decidable-first supervision: every supervisory question is first put to a deterministic gate that is reproducible by construction and returns an explicit UNDECIDED when the bytes do not contain the answer, so a model call is spent only on genuine residual ambiguity. Declared evidence contracts: each sub-task states up front what it will measure, and delivered evidence is diffed against the promise, which makes a silent sub-task a named failure rather than a success. Verified context delivery: a context block is accompanied by a verbatim needle checked against the assembled prompt, so ``delivered'' is a measurement rather than the producer's assertion. Promotion guards: a claimed metric is re-derived from the model's own out-of-fold output, fingerprinted to the evaluation set it was scored on, compared against a baseline ladder, and refused for degenerate or unprovable submissions. Around these sits per-agent instrumentation reporting what each agent cost, what it decided and whether its output reached a result, and a cross-run layer that refuses to compare runs whose evaluation sets or harness identities differ. We describe the architecture, the production measurements that motivated each mechanism, and what the design does not yet cover.