One Checkpoint, Three Robots, One Day: Zero-Shot Navigation as a System Property
Abstract
Zero-shot robot navigation is often evaluated as a property of the learned policy alone. In practice, however, successful deployment also depends on the platform components through which the policy perceives the environment, estimates motion, and executes actions. We study this interaction through two common navigation goal interfaces: a 2-D pose goal, which relies on metric odometry, and an image goal, which depends more strongly on the camera and lighting conditions. We deploy a single checkpoint, trained on 8.7 hours of teleoperation from one campus, unchanged and in a single day, on a Boston Dynamics Spot, an ANYmal D, and a 20 cm skid-steer rover. The rover’s camera is mounted four times lower than the lowest camera rig in the training corpus. We evaluate the same policy across courses ranging from a flat carpark to an unpaved bushland hillside, in daylight and at night. The transfer behaviour is determined more by the goal interface than by the robot or environment. Every uninterrupted pose-bearing run on every platform ended with an autonomous stop within 0.33–0.58 m of the goal. In contrast, image-only success on the same courses ranged from 0/5 to 5/5, depending on the camera and lighting conditions. The remaining deployment limits were set largely by the platforms rather than the policy, including compute, connectivity, and built-in protective layers. These results show that zero-shot transfer in the field is a system property, determined jointly by the policy, the platform around it, and the deployment environment.