Same Scene, Different Answer: Variance and Rollout Budgets for Flow-Policy Evaluation
Nihal Gunukula
Abstract
A reported success rate for a flow-matching vision-language-action policy is a sample from a distribution nobody has decomposed. We decompose it. Holding the initial state fixed leaves 61-84% of outcome variance in play across three policies on LIBERO: the scene accounts for the smaller share. A replicate block splits that remainder on the informative task into the flow sampler (38.5% of total) and system nondeterminism (46.0%), both invisible in a single reported rate; the partition replicates at three environment seeds. The consequence is a budget: rerunning the field's standard 500-episode protocol while changing only the flow-sampling seed spans 2.2 points, a band already containing 27% of the recomputable improvements highlighted across 50 recent VLA papers, of whose highlighted comparisons 34.4% state no usable rollout count at all. From the measured components we derive rollout budgets and the paired-design efficiency gain they imply, then spend them: a powered 4x3 factorial of denoising steps x execution horizon on released checkpoints shows the shipped default is dominated on $\pi_{0.5}$, where a single Euler step at horizon 10 is statistically indistinguishable from the 10-step default at 2.7x lower measured latency; $\pi_0$ does not tolerate the same cut. The cut's one confirmed price on $\pi_{0.5}$ is 4.7 points under camera-viewpoint shift, while replanning every step costs 9-41 points. Across four perturbation axes a 16.5-point clean gap stands for anything from 9.6 to 67 points of deployed margin, and rankings invert only where an axis drives the trailing policies toward the floor. Every analysis is pre-registered or labeled exploratory in a public deviations log; per-episode logs and the harness with an exact rollout-budget calculator accompany the release.
Chat is not available.
Successful Page Load