Evidence-Gated Bi-Level Autoresearch Across Models, Agents, and Systems
Abstract
Language-model agents can execute many experiments quickly, but reliable autoresearch also requires benchmark integrity, falsifiable plans, learning across branches, and stronger evidence before a system changes its own research policy. We present Svatah, a bi-level autoresearch system that separates project level discovery from evidence-gated improvement of the research process. In the inner loop, a Research Director compiles hypotheses into typed round plans; a critic admits benchmark-preserving experiments; isolated workers implement them; and an experiment DAG, role-aware memory, evidence packets, and decision trails carry verified knowledge forward. In the outer loop, recurring research failures can trigger bounded changes to a typed system policy, with replay, development, held-out, cross-provider, and reversible-canary gates before activation. Three cross-domain campaigns demonstrate the inner loop’s range. Qwen3 1.7B GSM8K accuracy increased from 12.36% to 45.94% (3.72×); an exact 100-million-element GPU kernel improved from 31.07ms to 3.087ms (10.06×); and agentic-small-language-model validation success increased from 4.35% to 30.43%. In frozen supporting evaluations, Svatah configurations improved multi-objective official reward by up to 55.7% over the neutral baseline and reduced solar hidden MAE by 6.3%. Across RSI-enabled runs, 16 policy candidates were evaluated and the evidence gates preserved the active policy when transfer criteria were not met. These results position evidence-gated meta-agent orchestration as a practical architecture for auditable research across model training, agent adaptation, and systems optimization.