Adjoint Guidance Flow for Efficient Critic-Guided Vision-Language-Action Models
Abstract
Flow-based Vision-Language-Action (VLA) policies are typically trained with behavior cloning and thus do not explicitly optimize long-term task return. Critic-guided methods address this by steering action generation with value gradients, but require costly critic evaluation and back-propagation at every generation step. We propose Adjoint Guidance Flow (AGF), a lightweight guidance network that amortizes critic-based guidance for flow-based VLA policies. The network is trained with adjoint matching, which propagates terminal critic gradients through frozen flow dynamics to provide trajectory-aware supervision at intermediate action states. At inference time, AGF directly predicts value-informed corrections with a single forward pass, eliminating the need for critic ensembles, critic back-propagation, and adjoint computation. AGF improves task success on LIBERO and RoboCasa while running up to 15.1x faster with 17.7x fewer trainable parameters than prior critic-guidance methods, all while preserving the pretrained VLA policy.