A Process-Level Evaluation of LLM Discovery Agents
Haorui Wang ⋅ Yuanqi Du ⋅ Hao Cheng ⋅ Baolin Peng ⋅ Chao Zhang ⋅ Weizhu Chen ⋅ yelong shen
Abstract
Large language models (LLMs) are increasingly used as \emph{discovery agents} that iteratively edit codes against evaluators on tasks from mathematical optimization, algorithmic problem-solving and scientific discovery. Nevertheless, current evaluation reports only a single final score per (model, task) pair under a fixed agent harness: the surrounding scaffolding that governs feedback, memory, reflection, and revision. This final-score view obscures how agents search, where they fail, and which harness mechanisms are actually useful for different classes of tasks. In this paper, we introduce a four-level ladder of agent harnesses and a process-level evaluation framework that measures behavioral profiles and intermediate milestones, alongside the final performance. Using the evaluation workflow, we run a fully crossed evaluation of six contemporary LLMs on $24$ curated discovery tasks across four harness configurations, producing $1{,}728$ trajectories. Our analysis reveals two notable findings. First, harness effectiveness is highly problem-dependent. Second, models can exhibit distinct search behaviors even when their final scores are similar. We release all the trajectories, behavioral profiles, capability rubrics, and harness implementations as reusable evaluation resources for the community.
Chat is not available.
Successful Page Load