How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
Abstract
Large language models and agentic scaffolds are reshaping scientific research: a single system can now carry a study from an initial hypothesis through to a written paper—a paradigm now referred to as AutoResearch. Existing evaluations of these agents largely score final outcomes, revealing little about what agents actually do during a run or why apparently successful trajectories fail scientifically. We introduce AutoResearchEval, a corpus of 800 complete autonomous research-agent trajectories from eight harness–model combinations on 100 real-world frontier research tasks spanning seven scientific domains and the full research lifecycle; each trajectory retains the agent’s planning, retrieval, tool use, code, generated data, intermediate artifacts, report, and review. From grounded-theory analysis of expert-annotated trajectories we induce ARFT (AutoResearch Failure Taxonomy), 45 behaviorally grounded failure patterns on two cross-cutting axes, and to apply it at scale we develop a human-calibrated, artifact-aware agent-as-a-judge that detects failures invisible in endpoint metrics or final reports alone. We find that agents frequently produce valid-looking outputs through metric substitution, circular validation, unsupported claims, and self-diagnosed but uncorrected flaws. While these behaviors span all stages of the lifecycle, they converge on a single overarching limitation: current agents lack a metacognitive loop—the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound. The same patterns recur across all eight combinations, including the strongest tested, indicating a general limitation rather than an artifact of any one scaffold; whether orchestration-level interventions can close it is an open question this work does not test. We publicly release AutoResearchEval and ARFT as a dataset, a vocabulary, and a scalable method for interpreting long-horizon agent behavior.