OpenMLE: Training Evolutionary MLE Agents toward Recursive Self-Improvement
Abstract
Using AI to build and improve AI requires agents that can propose algorithms, execute experiments, diagnose failures, and decide how to spend the next unit of compute; machine learning engineering (MLE) is a concrete, executable testbed for that loop. We present OpenMLE, an open full-stack solution for training language models to construct and iteratively improve machine learning solutions from execution feedback. OpenMLE-Gym turns curated anchors, Kaggle datasets, and Kaggle competitions into 5,758 quality-gated executable task packages with isolated execution and task-specific evaluators; OpenMLE-ERL converts verified solutions and revisions into execution-grounded supervision and reinforcement learning for four reusable program-transformation operators; and OpenMLE-Evo composes the same operators into long-horizon search using structured experience, non-greedy parent selection, and operator-conditioned memory. Because one operator interface is shared by post-training and inference, verified search trajectories train exactly the transformations the harness later composes: the improver is trained, while the harness, evaluator, and stop policy stay fixed. Under an identical harness, OpenMLE-35B raises MLE-Bench Lite Medal Average from 39.39% to 60.61% and Human Rank from 0.5828 to 0.7647, with OpenMLE-30B reproducing the direction on a second backbone; matched harness comparisons favour OpenMLE-Evo over general-purpose scaffolds and over original AIRA-Evo; and OpenMLE-35B reaches 71.21% with OpenMLE-Evo-Max, which adds benchmark-disjoint experience priors and asynchronous parallel search. Controlled NatureBench Lite comparisons indicate that both transfer beyond competition-style MLE. We read this as one concrete step from evolution toward recursive self-improvement, not a realization of it.