Beyond Accuracy: A Diagnostic Benchmark for Hypothesis-Driven Experiment Planning in LLM Agents
Abstract
Evaluating large language model (LLM) agents on scientific tasks typically reduces to final-task accuracy, which obscures where reasoning succeeds or breaks down. We introduce a benchmark and diagnostic evaluation framework for hypothesis-driven adaptive experimental planning, instantiated in kinetic mechanism identification. The framework goes beyond final accuracy in two ways. First, we evaluate exploration behavior by quantifying the extent to which agents exhibit space-filling or corner-seeking patterns in the experimental design space. Second, it decomposes the reasoning process into four interpretable steps; belief update, hypothesis discrimination, discriminative experiment design, and predictive reasoning under interventions, each scored independently on reasoning traces. Across ten LLM agents, we find that stronger agents outperform both adaptive (Bayesian optimization) and non-adaptive (Latin hypercube) baselines and exhibit exploration patterns inconsistent with naive heuristics. Decomposing reasoning reveals that the stronger agents score consistently well across all four steps, whereas weaker agents often succeed at belief update and hypothesis discrimination but fail at discriminative experiment design and predictive reasoning. This contrast shows that our framework identifies a specific reasoning bottleneck: weaker models can recognize plausible hypotheses and articulate their differences, but struggle to translate this understanding into concrete predictions and experimental designs. The code and dataset are available at https://github.com/krfdq48k4p-arch/HypothesisDrivenExperimentPlanningBench and https://huggingface.co/datasets/xq8wvm/HypothesisDrivenExperimentPlanningBench, respectively.