Simple Baselines are Competitive with Code Evolution
Abstract
Code evolution is a family of techniques that rely on large language models to search through possible computer programs by evolving existing code. While code evolution pipelines have shown impressive performance across domains, many are highly complex and are typically not compared to simpler alternatives. To prevent bad comparisons from stagnating the field, we propose two simple baselines and compare them to popular pipelines across three domains: finding better mathematical bounds, designing agentic scaffolds, and machine learning competitions. Surprisingly, the baselines match or outperform existing methods across all domains, prompting a closer look at which factors drive performance in different settings. For mathematical bounds, we find that the search space matters far more than the exact search algorithm, with both simple and complex methods performing similarly given the same search space. Thus, expanding the search space is what matters for improving performance, and does so across all pipelines. Moreover, different domain knowledge embedded in prompts can significantly affect the search's efficiency, potentially leading to overestimating a pipeline's sample efficiency. For automated agentic scaffold design, we show that typical high-variance evaluations lead to all automated search methods picking subpar scaffolds relative to a simple majority vote. To mitigate this, we propose improved evaluation procedures that reduce stochasticity while keeping the search economically feasible. Overall, these results indicate that further improving code evolution's performance depends more on the setup surrounding the search pipeline than on the pipeline itself. We hope these insights will enable developing new code evolution methods and help spur meaningful progress in the field.