AdaMAP: Learning Adaptive Multi-Action Prediction with Grounded Dreaming Guidance
Abstract
Large agentic models can now operate in complex digital games by pretraining on web-scale data, but gameplay trajectories often contain highly redundant actions. Single-action agents can exploit this redundancy spuriously by copying recent actions, leading to causal confusion. Fixed-length action chunks reduce this ambiguity by predicting multiple future actions at once, but their open-loop execution accumulates errors because the agent cannot adapt to new observations during the chunk. The agent should therefore learn when to act for longer without feedback and when to stop early for a new observation. We introduce AdaMAP (Adaptive Multi-Action Prediction), an adaptive multi-action prediction framework that directly generates variable-horizon action chunks, interleaves actions with grounded visual dreams that provide verifiable future constraints, and optimizes the resulting policy with offline imitation learning followed by online MAP-PPO. Across Minecraft and VizDoom, AdaMAP attains the best or tied-best result in five of six task categories and improves average Minecraft success by 3.56 percentage points over the fixed-horizon RL baseline. The learned horizons vary systematically across tasks and games, suggesting that adaptive multi-action prediction discovers useful temporal abstractions.