Denoising-Time Heterogeneity in VLA Action Generation: A Controlled Study via Step-Wise Expert
Abstract
Denoising-based Vision-Language-Action (VLA) policies are often parameterized by a single denoiser shared across denoising time. This paper shows that expert usefulness in VLA action generation can vary systematically across denoising phases, and that this variation can be exposed through controlled step-wise expert aggregation. We mix two full action experts at the velocity-field level with a fixed smooth schedule, enabling schedule-only interventions that keep expert capacity and total mixing mass matched while changing when each expert is emphasized. The empirical evidence is organized as a matched chain: semantically distinct expert pairs test denoising-phase assignment through forward versus reversed schedules, while an anchor-free isomorphic-expert control tests schedule sensitivity without manually designed expert semantics. Experiments on both LIBERO and CALVIN show the same directionality phenomenon: aligned schedules outperform their time-reversed counterparts under matched strength, indicating that when each expert is emphasized matters for action denoising. We complement these findings with a squared-error risk analysis showing that denoising-time variation in expert advantage implies a time-dependent optimal mixture, together with a schedule-weighted objective note explaining why copied experts can still functionally differentiate under non-constant schedules.