DualVLN on Jetson Thor: Measuring Inference Cost, Observation Freshness, and Navigation Quality
Abstract
A vision-and-language navigation (VLN) model on a robot must fit the onboard computer and act on observations that are still current. We port DualVLN, which pairs a 7B planner (System 2, S2) with a diffusion trajectory model (System 1, S1), to a single Jetson Thor without retraining and evaluate H1 humanoid navigation in physics simulation. Against the original PyTorch execution of the public checkpoint on the same device, our TensorRT port with an NVFP4 planner cuts median latency by about 6.8x for planning, 7.9x for look-down planning, and 3.9x for trajectory generation, timed at different boundaries. A separate replay shows 58.5% lower peak RAM above idle. On the 12 paths used for selection, the port lowers the success rate from 0.556 to 0.458. Because it also changes engines and process configuration, the drop cannot be attributed to NVFP4 alone, although NVFP4 may cost quality. Among four runtime settings, using each S1 output once, not for up to four actions, lowers the 95th-percentile age of the current observation from 2.91 to 0.99 s at a similar S1 service rate. The final setting, selected and evaluated on the same 12 paths, adds four successes there over the port's default setting, all from three paths, and ties the setting before it. On 12 other paths, the gain does not recur and the drop neither repeats nor is ruled out: the original execution, the port, and the final setting succeed in 20, 23, and 22 of 48 runs each. In this small simulation study, where the scene is held while the models compute, inference speed alone determines neither observation freshness nor navigation quality.