From Ideas to Code: Tree-structured Policy Optimization for Automated Algorithm Design with LLMs
Abstract
Large Language Model-based Automated Algorithm Design (LLM-AAD) has shown promising results across diverse domains by coupling LLMs with iterative search. Some recent efforts have begun to fine-tune LLMs for algorithm design through reinforcement learning. However, these attempts still rely on outcome-level supervision, leaving intermediate ideas and plans without direct credit, even though they determine the conceptual direction of algorithm design. In this paper, we propose Algorithm Tree Policy Optimization (ATPO), which utilizes tree-structured rollouts to decouple the generation of conceptual ideas from their specific code implementations. By applying group-relative advantage estimation, ATPO assigns credit to both strategic ideas and algorithmic implementations, thereby optimizing the policy across multiple levels of abstraction. Integration with FunSearch, EoH, and OpenEvolve demonstrates improvements over their original counterparts. Notably, the learned policy transfers effectively to unseen problem instances and alternative search methods, indicating that it can improve the generalization of algorithm design policies.