OpenMLE: Training Evolutionary MLE Agents toward Recursive Self-Improvement
Abstract
Long-horizon agents need an execution layer that coordinates models, tools, state, resources, and verification without sacrificing reproducibility. We present OpenMLE, an open full-stack solution for training language models to construct and iteratively improve machine learning solutions from execution feedback, and read it as a concrete agentic-OS design. OpenMLE-Gym turns curated anchors, Kaggle datasets, and Kaggle competitions into 5,758 quality-gated executable task contracts served by a centralized scheduler over isolated CPU/GPU workers that persist logs, submissions, scores, and artifacts. OpenMLE-ERL converts the records that runtime already writes into execution-grounded supervision and reinforcement learning for four reusable program-transformation operators, with asynchronous rollout consumption preventing heterogeneous program runtimes from stalling policy updates. OpenMLE-Evo keeps a structured experience board with lineage, operation type, and validation progress, and composes the same operators into bounded long-horizon search. Because one execution contract is shared by task construction, post-training, and inference-time search, a model swap and a runtime swap are separately measurable: under a fixed runtime OpenMLE-35B raises MLE-Bench Lite Medal Average from 39.39% to 60.61%, under a fixed checkpoint the runtime raises 53.03% to 60.61%, and OpenMLE-35B reaches 71.21% with OpenMLE-Evo-Max -- benchmark-disjoint experience priors and asynchronous parallel search -- while spending 41.7% fewer model tokens per run. We close with the interfaces, replay semantics, and benchmark contract such a layer still needs.