Counterfactual Instruction Grounding for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models inherit rich semantic priors from pretrained vision-language models, yet these priors do not guarantee persistent instruction sensitivity after narrow-domain fine-tuning. We observe a systematic failure mode: although VLA models remain visually reactive, their later actions can become weakly constrained by the commanded instruction, leading to unstable object interaction and visually plausible but instruction-inconsistent behavior. We trace this behavior to a collapse of instruction-conditioned representations across network depth and task progress. More fundamentally, this collapse is enabled by the learning objective itself: standard imitation learning admits Bayes-optimal solutions that are invariant to language. We address this by recasting instruction grounding as identifying the commanded instruction among counterfactual alternatives. Based on this view, we propose Counterfactual Instruction Grounding (CIG), a contrastive objective that encourages the generated trajectory to remain identifiable with the commanded instruction among counterfactual alternatives. CIG applies to both autoregressive and flow-matching VLA models, using exact action-chunk likelihoods for autoregressive models and an energy-based pseudo posterior for flow-matching models to avoid intractable trajectory likelihood estimation. Practically, CIG can be applied directly to already fine-tuned models as a continuation-stage grounding objective, restoring instruction sensitivity without retraining from scratch or changing the architecture. Extensive experiments on RoboCasa and LIBERO-Plus demonstrate that CIG achieves strong instruction-following performance with success rates of 71.3% and 86.3%, respectively. Real-world deployment further shows that CIG transfers to kitchen manipulation and preserves instruction-conditioned behavior in visually ambiguous scenes.