AutoWALNUTS: Coverage-First Geometry for Exact HMC via Guarded Autoresearch
Victor Gallego
Abstract
Autonomous coding agents can explore algorithmic ideas quickly, but noisy benchmarks make it easy to overfit a target suite or exploit quirks in the evaluator. We study how to prevent these failures when developing Markov chain Monte Carlo (MCMC) methods. Our framework, guarded autoresearch, gives agents a typed target–sampler interface while keeping evaluator assets immutable and separating development, validation, and held-out targets. Paired seeds and distributional tests provide empirical checks, while explicit exactness conditions ensure that the final kernel preserves the target distribution. Across four autonomous runs, agents carried out more than 100 experiments and found complementary improvements. We combine them in AutoWALNUTS. Its warmup first explores the unconstrained coordinates, then learns a guarded nonlinear scale chart and dense whitening map. After freezing that geometry, it samples with a fixed WALNUTS transition in pullback coordinates. On nine held-out targets, AutoWALNUTS improves minimum ESS per gradient by roughly $6$-$13\times$ over walnutpie, the reference C++ WALNUTS implementation. These results suggest that AI agents can help develop reliable scientific software when evaluation is staged, information flow is controlled, and domain requirements are built into the search. Code at https://anonymous.4open.science/r/autowalnuts
Chat is not available.
Successful Page Load