MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
Abstract
Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, multi-step navigation, and future state changes. Yet many agents externalize this computation as long textual thoughts, making interaction slower, supervision more costly, and deployment less efficient. We introduce MIRAGE (Mobile agents with Implicit Reasoning And Generative world modEls), a framework that learns continuous latent reasoning representations from visible textual thoughts. MIRAGE introduces an efficient latent-space learning procedure that transfers explicit reasoning into compact hidden states, allowing the agent to reason internally without decoding long rationales. It further brings a world-model perspective into mobile-agent training: the model’s latent reasoning vectors are aligned with future screenshots, encouraging the agent to predict upcoming interface states in latent space before executing an action. This makes the hidden computation not only a compressed thought trace, but also a forward-looking representation of how the environment may change. At inference time, MIRAGE reasons in continuous latent space, reducing token generation while improving execution efficiency. On AndroidWorld, MIRAGE matches explicit-CoT SFT in the 4B ablation under a 3–5× lower decoded-token budget, and improves a comparable instruction-tuned baseline by 10.2 points; on AndroidControl, it improves action grounding with over 75% fewer generated tokens.