3A-VLA: Abstraction-Aligned Action Learning for Vision-Language Agents in 3D Game Worlds
Abstract
Despite advances in vision-language-action (VLA) models, they remain limited in highly dynamic game settings such as 3D open worlds and competitive player versus player (PvP) settings, where agents must efficiently extract sparse, actionable signals from dense visual input while integrating multimodal cues to track off-screen high-value targets in real time. To address this, we introduce 3A-VLA (Abstraction-Aligned Action VLA), a framework that grounds action learning in explicit intention and environment abstractions rather than superficial pattern matching. We introduce dual task-agnostic abstractions: the intention abstraction (IA), which condenses verbose instructions and reasoning into explicit semantic primitives, and the environment semantics abstraction (ESA), which structures dense visual streams into a spatial--functional affordance representation to guide grounded actions. We further propose an abstraction alignment reweighting (AAR) module that adaptively reweights the action imitation loss. A continuous intention--environment alignment signal is used to emphasize reliable action supervision when abstractions agree and reduce the influence of ambiguous demonstrations when they diverge, thereby learning a policy that balances fine-grained control and high-level reasoning without the need for manual rules. Extensive experiments show that 3A-VLA yields state-of-the-art results in both open-world (Minecraft) and competitive PvP (Game for Peace) settings. It also demonstrates strong zero-shot generalizability to high-fidelity games across different domains, including Valorant, CS2, GTA V, Elden Ring, and Mount & Blade. Code and models will be publicly released at https://3a-vla.github.io/.