FactoryBench: Evaluating Industrial Machine Understanding
Abstract
Robot foundation models are increasingly deployed zero-shot: downloaded pretrained and pointed at a machine they were never trained on. We ask what such a model understands about a physical system it is seeing for the first time, and introduce FactoryBench, a zero-shot benchmark over dense industrial robot telemetry. Q&A pairs span four levels (state estimation, intervention, counterfactual reasoning, and engineering decision-making) and five answer formats, four scored deterministically and free-form answers by an LLM-as-judge protocol. We release FactoryWave (a dense multitask sensor dataset from a UR3 cobot and a KUKA KR10 industrial arm) and build FactoryBench as over 70k items grounded in roughly 15k episodes from FactoryWave, AURSAD, and voraus-AD. Across six frontier models, none exceeds 50% (chance-corrected) on the structured levels, two fail to beat a linear-regression baseline on state estimation, and the ranking reshuffles entirely on decision-making, with the leader on the first three levels falling to the bottom of the fourth. A human expert reaches 0.96-0.98 on the same items, so the headroom is real rather than an artifact of unanswerable questions. Zero-shot transfer to an unseen industrial machine is not yet reliable enough to act on.