Heuresis: Evaluating Search Strategies for Autonomous Machine Learning Research Agents
Abstract
Autonomous AI research promises to accelerate the scientific progress of machine learning. While current Large Language Model (LLM)-based agents excel at writing code, the bottleneck in research is the exploration of diverse and novel ideas: agents collapse to common techniques present in their pretraining data, or prematurely converge on suboptimal solutions (Padmakumar et al., 2024; Jiang et al., 2025). To this end, we introduce Heuresis, a framework that abstracts the research pipeline into a set of general and composable primitives, enabling open-ended scientific exploration in machine learning research. We implement 5 known algorithms spanning Quality-Diversity, Evolutionary, and Curiosity-based search, in addition to a greedy baseline, and evaluate them across three axes -- Quality, Diversity, and Novelty -- on two domains: LLM pretraining and On-Policy RL. We find that quality and novelty are inversely correlated. Methods that optimize purely for raw performance reach the strongest single solutions on a given task by replicating prior work: the top-quality runs from the greedy baseline are uniformly classified as direct copies under the Gupta-Pruthi rubric (Gupta & Pruthi, 2025). The few high-quality, verified-novel ideas in our study come from algorithms that balance performance with diversity or curiosity-based exploration. We also observed that agents resorted to a variety of reward-hacking techniques during execution, whose detection was necessary to keep the search faithful to the task. Our results underscore the importance of search and Quality-Diversity methods for autonomous research, and our framework opens up opportunities for further inquiry towards the ultimate goal of perpetual, autonomous scientific progress.