OT-Robust3DVLA: A Wasserstein Barycenter is the Right Inductive Bias for Robust Multi-View 3D Vision-Language-Action Policies
Pawan Kumar
Abstract
Multi-view 3D vision-language-action (VLA) policies are brittle to small sensor perturbations: depth jitter, partial occlusion, calibration drift and per-view dropout can move the predicted action chunk by far more than the perturbation itself. We trace the brittleness to the *fusion stage*: concatenating per-view tokens lets a small geometric shift flip arbitrary token indices, turning a continuous-mass-transport problem into a discrete indexing problem. We propose **a single substitution**: replace concat-then-attention fusion with a *differentiable Wasserstein barycentric tokenizer* that fuses $K$ view measures into $M$ barycentric tokens by transport. We prove a chain of Lipschitz-style stability bounds that hold conditional on a declared OT perturbation set, calibrated Lipschitz constants, and bounded Sinkhorn-solver residuals. Empirically, the barycenter is the single dominant mechanism behind the gain: (i) the proposed model has the lowest mean clean $\ell_2$ on synthetic, LIBERO, RLBench and CALVIN ABC→D (3-of-4 with non-overlapping 95% CIs); (ii) a controlled ablation on synthetic regresses by $46\times$ on the corruption-stability metric when the barycenter is removed and by $<1.5\times$ when any other component is; (iii) the same mechanism transfers to LIBERO with a $+37\%$ clean $\ell_2$ regression and $+62\%$ action-OT regression on real data; (iv) in synthetic closed-loop the concat-fusion baseline collapses from 85% clean to 0% under per-view dropout while the barycenter keeps a 92% worst-case across all five conditions (3 seeds, narrowest CI of all four methods).
Chat is not available.
Successful Page Load