Robot Foundation Models in the Materials Lab: Benchmarks, Embodiments, and Open Challenges
Abstract
Robot foundation models, exemplified by vision--language--action (VLA) policies, now execute diverse language-specified manipulation in everyday environments and are beginning to be evaluated in chemistry and materials laboratories. This paper presents a systematic survey of this emerging intersection, reviewing recently released laboratory benchmarks, simulators, and real-robot systems along four axes, namely which experimental domains are covered, which embodiments and hands are used, which base models and data strategies are employed, and how performance is evaluated. The evidence shows a consistent pattern in which policies that succeed at everyday tasks fall short of laboratory requirements, with task completion systematically overstating experimental validity in precision, safety, and protocol compliance. We distill seven open challenges and outline possible directions for a materials-native research agenda.