PHYSICIAN: Verification of Physical Grounding in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have demonstrated strong semantic reasoning for embodied tasks, but their action predictions can remain physically ungrounded. In particular, a command that is semantically appropriate may be infeasible under the actual mass, friction, contact, and dynamic constraints of the environment. We characterize this failure as a Physical Grounding Gap, in which the physical assumptions underlying a VLA action diverge from the dynamics governing its execution. This gap can produce kinematic and dynamic hallucinations that are difficult to detect from visual-language reasoning alone. We propose PHYSICIAN, a generative middleware safety kernel for detecting and mitigating physical-grounding failures in VLA-controlled robotic systems. PHYSICIAN implements a four-layer Verification-in-the-Loop architecture that combines learned visual reasoning, deterministic physical simulation, generative counterfactual prediction, and runtime safety enforcement. An agentic vision module, powered by Gemini 2.5 Flash, performs Think-Act-Observe inference over visual telemetry to estimate latent physical attributes, including friction coefficients and object mass. These estimates are used to construct two independent predictions of the consequences of a candidate action: a deterministic prediction generated through PyBullet simulation and a generative prediction produced by a video diffusion world model. We use the divergence between these predicted futures as an empirical uncertainty signal for identifying actions whose physical consequences are insufficiently grounded. Actions associated with high uncertainty or predicted constraint violations are subsequently filtered by a Control Barrier Function (CBF) layer before execution. If verification thresholds are exceeded, PHYSICIAN triggers an emergency interrupt and transitions the system to a predefined safe state. We evaluate the framework in multi-degree-of-freedom, contact-rich manipulation environments containing deceptive low-friction surface transitions and dynamic payload uncertainty, and further validate its operation through sim-to-real deployment on an AlphaX 4WD mobile telemetry platform. Under deceptive low-friction conditions, task success improves from 12% to 89%, while performance under dynamic payload uncertainty improves from 38% to 82%. The asynchronous dual-path architecture performs end-to-end verification in under 550 ms, maintaining compatibility with real-time robotic control. These results indicate that combining learned generative predictions with deterministic physical verification can provide an effective mechanism for identifying physical-grounding failures in VLA models and improving their reliability under uncertain and previously unseen physical conditions.