Phase-Adaptive Fusion: Spatio-Temporal Modulation for VLA Models
Abstract
Vision-Language-Action (VLA) foundation models have driven remarkable progress in general-purpose robotic manipulation. Yet, adapting these models to dynamic physical environments exposes a severe vulnerability: a modality imbalance during optimization. Because low-dimensional proprioceptive states provide a more direct path for loss minimization, policies instinctively form a "state-dominant shortcut," systematically suppressing the learning of high-dimensional visual features. This visual neglect causes models to bypass precise spatial grounding, leading to catastrophic alignment failures during contact-critical execution phases. To rectify this without disrupting natural training dynamics or requiring heavy auxiliary decoders, we introduce Phase-Adaptive Fusion (PAF), a lightweight, non-intrusive neural modulation plug-in. Operating entirely in the forward pass, PAF features a decoupled dual-branch architecture. A spatial branch performs explicit semantic-geometric alignment, while a temporal branch infers latent task execution phases from recent action history and state kinematics. By synthesizing these signals into a residual gating mechanism prior to multimodal fusion, PAF dynamically amplifies visual tokens precisely during alignment-heavy stages, all while preserving the host VLA's pre-trained representational manifold. Extensive evaluations across simulated benchmarks (LIBERO, Meta-World, Robomimic) and real-world long-horizon tasks demonstrate that PAF successfully breaks the proprioceptive shortcut, drastically reducing terminal alignment drift and significantly enhancing overall manipulation robustness.