BEAKER: An Expert-Curated Benchmark for Embodied Brains in Self-Driving Chemical Laboratories
Abstract
Self-Driving Chemical Laboratories (SDCLs) are moving chemical experimentation toward embodied automation, where machines must perceive laboratory scenes, reason about experimental states, and act under physical, procedural, and safety constraints. Although Multimodal Large Language Models (MLLMs) are increasingly considered as embodied brains for such systems, existing laboratory benchmarks mainly focus on safety, anomaly detection, or general scientific understanding, leaving their execution-oriented embodied capabilities underexplored. To fill this gap, we introduce BEAKER, a Benchmark for evaluating Embodied Actionability, Knowledge, Experimentation, and Reasoning in SDCLs. BEAKER contains 1,000 expert-curated QA samples across 11 image and video task types, covering key judgments required for chemical laboratory execution, including affordance grounding, trajectory planning, equipment and state understanding, apparatus connection, operation process reasoning, and result understanding. All questions and annotations are manually created and reviewed by chemistry and embodied AI experts, without AI-assisted generation or annotation. We evaluate 28 general-purpose, scientific, chemical, and embodied MLLMs on BEAKER. Results show that current models can recognize laboratory objects and states, but still struggle to ground experimental knowledge into operable regions, feasible trajectories, and long-horizon procedural reasoning. BEAKER thus reveals a clear gap between multimodal laboratory perception and action-grounded experimental understanding, providing a diagnostic benchmark for future MLLMs in SDCLs.