PostTrainBench$^{0}$: Can LLM Agents Automate LLM Post-Training Without Gradients?
Yuqiao Tan ⋅ Minzheng Wang ⋅ Shizhu He ⋅ Jun Zhao ⋅ Kang Liu
Abstract
Developing a post-training strategy requires choosing experiments, interpreting feedback, and revising search procedures. We introduce PostTrainBench$^{0}$, a benchmark that places these decisions under the control of a large language model (LLM) agent. Given a base model, editable search code, and a four-hour budget, an agent designs a search program that generates, combines, and refines seeded weight perturbations. A fixed evaluator provides search feedback for candidate selection; the selected checkpoint is assessed on held-out examples spanning mathematics, coding, narrative reasoning, and chemistry. We evaluate ten LLMs with multiple reasoning-effort settings on Qwen2.5-3B-Instruct and Qwen3-4B-Base. Tuned evolution strategy (ES) attains the highest mean held-out score on both targets, establishing a reference for the agents' use of the supplied search methods. Task-level results reveal shared capability bottlenecks, with much of the ES advantage concentrated in a few tasks. Execution traces show how agents restart search, change update rules, and revise direction combinations in response to feedback. Comparisons of search progress further distinguish agents with similar early scores and evaluation volumes. By combining editable search programs, a shared evaluation protocol, and replayable checkpoint definitions, PostTrainBench$^{0}$ connects agents' decisions to search progress and final model capabilities. It makes the productive use of algorithms and feedback a concrete target for evaluating autonomous optimization research.
Chat is not available.
Successful Page Load