H2L-Bench: A Benchmark for Agentic Planning Under Uncertainty in Drug Discovery
Pablo Lemos ⋅ Nhung T Nguyen ⋅ Kevin Kühl ⋅ Aditya Kashyap ⋅ Ahmed Elnaggar ⋅ Dong Ding ⋅ Majd Mustapha ⋅ Umut Eser ⋅ Timothy Schultz
Abstract
AI agents have made rapid progress on tool use, multi-step planning and open-ended reasoning, yet the benchmarks that drive this progress rarely capture what matters in scientific decision making: acting under uncertainty, choosing between costly experiments, and committing to an answer with limited information. We introduce H2L-Bench, a lightweight decision-theoretic sandbox that simulates the hit-to-lead phase of small-molecule drug discovery as a budget-constrained sequential-decision task with a mechanistic ground truth. An agent chooses among eight assays of varying cost and information content, then nominates one compound from a library. Winning requires clearing multiple hidden thresholds simultaneously, with costs and noise scaled across three difficulty tiers. As a principled reference we implement a Bayesian value-of-information (VoI) agent that maintains an analytical Gaussian posterior over each compound's latent traits and picks actions by decision-theoretic expected information gain per dollar. Comparing VoI to a greedy baseline and three frontier Claude LLMs (Haiku 4.5, Sonnet 4.5, Opus 4.5) at 50 seeds per cell, we find a consistent ordering (greedy $<$ Haiku $<$ Sonnet $\approx$ Opus $<$ VoI) with the largest gap on easy tier (VoI 88\% vs Opus 63\%) and a genuine tie on medium (54\% vs 47\%). Beyond the headline win rates, the environment lets us watch where LLMs and a Bayesian agent diverge action-by-action: on a subset of seeds, LLMs recover the correct compound using two-turn patterns, or taking a risky bet and committing budget to an expensive singular assay, that a depth-1 per-dollar VoI ranking cannot propose.
Chat is not available.
Successful Page Load