Conservative Meta-Agent Search: Architecture Adaptation under Non-Stationarity
Ali A Alzahrani
Abstract
Automated design of agentic systems is typically evaluated on static task suites with immediate scalar rewards. Deployed systems face persistent shift, delayed dependent feedback, correlated component failures, and costly architecture changes. We formalize meta-agent architecture search as sequential incumbent replacement with switching costs and hard risk constraints, and introduce Conservative Meta-Agent Search (CMAS). Throughout, meta-agent denotes the entire outer loop: proposer, compiler, evaluator, gate, and rollback controller. CMAS retrieves archived architectures and proposes typed graph edits, then scores each against the incumbent by paired shadow replay under common random numbers. Component credit is interventional. Promotion requires an empirically calibrated, variance-penalized simultaneous score above a pre-specified margin, and is followed by a simulated canary stage and pre-specified rollback. Backtested return cannot certify meta-level skill, so we introduce MetaInvest-Bench, a regime-switching simulator whose specified worker competence makes a surrogate-optimal architecture computable, with leakage-controlled replay, adversarial governance tracks, and five reusable negative controls. Financial streams give an instrumented, risk-sensitive test of meta-level generalization; we claim no alpha. Across 96 streams, CMAS cuts normalized mean surrogate dynamic-oracle regret from $0.116$ to $0.104$ and raises held-out utility from $0.44$ to $0.48$ over the strongest adaptive common-budget reimplementation; when extra workers add no value, it contracts to near-single-agent teams. Its deployment inference cost is $\sim$6% lower, though its search-inclusive token cost is the highest of the methods compared. The observed per-update false-promotion rate is 3.4% against a nominal 5% target, and 4.8% in a sealed rerun. The statistical ingredients are standard. The calibration is empirical: it concerns an update-level replay statistic on an author-designed benchmark, and guarantees neither full-utility improvement, nor absence of simulated canary-stage harm, nor sequence-level error control.
Chat is not available.
Successful Page Load