Decoupling Physical Safety from Robot Foundation Models: A Review and Case Study
Abstract
As robot foundation models and Vision-Language-Action (VLA) architectures are increasingly deployed for embodied manipulation, ensuring physical safety remains an open challenge. Only recently have approaches begun to ground physical safety in spatial recognition, using these models understanding of scenes, object boundaries, and workspace layout to determine where an agent can safely act. In this paper, we review this emerging link between spatial recognition and physical safety in embodied foundation models, examining how different methods recognize and represent safety-relevant spatial structure, and what each approach offers and lacks. We then present a decoupled approach to enforcing physical safety, via a case study that demonstrates how language models can perform spatial recognition of a scene to generate geometric bounded safety regions, enforced independently using a Python Operational Space Control Barrier Function (OSCBF) package. This case study is evaluated using seven small open-source language models in a PyBullet simulation environment, and we show that models prompted with a structured prompt produced minimally sized task-oriented safety regions across metrics for task performance and collision avoidance. This paper concludes with open research directions for grounding physical safety in the spatial recognition capabilities of embodied foundation models.