Evaluating Physical Reasoning in LLM Agents Requires Construction Benchmarks
Abstract
Physical reasoning is central to agents that step into the physical world. Models now excel at code and math, but their physical reasoning capabilities remain largely untested. Existing physical benchmarks evaluate goal-state matching rather than functional performance. This position paper argues that evaluating physical reasoning in LLM agents requires construction benchmarks, where physics verifies functional outcomes. To operationalize physical reasoning evaluation, we identify five capabilities spanning the full agentic loop from goal to verified artifact. Our audit of eleven physical benchmarks finds only two that combine realistic physics with functional verification, the minimum requirements for testing physical reasoning; both are construction benchmarks. Neither covers all five capabilities, yet frontier models already achieve near-zero success as task complexity increases. Progress depends on the breadth of benchmarks we build. We call on academia, game studios, and industry labs to expand construction benchmarks along both axes, widening capability coverage and deepening physical challenges. Their breadth determines how fully agents move from code and math to reshaping the physical world.