AndroidFlux: Evaluating and Reward Modeling Failure Recovery Capabilities in Mobile Use Multimodal Agents
Abstract
Mobile-use multimodal agents are judged almost entirely by terminal task success from a clean start. That measure is saturating and, more fundamentally, too sparse to evaluate intermediate failure and recovery: a single binary outcome per episode cannot separate an agent that derails at the first unexpected error from one that recovers and finishes successfully. Recovery matters both for robustness in the wild and for the multi-agent frameworks now used to control inference cost, which route work across a pool of agents and escalate episodes a weaker agent fails to a stronger one. The resuming agent inherits a foreign context and a device state that appears in no clean demonstration, so whether it can recover determines whether such systems are viable — yet no mobile-use benchmark measures it. We therefore introduce AndroidFlux, a failure-recovery benchmark built on AndroidWorld. Recorded trajectories are replayed to controlled checkpoints around the first recoverable error, and a target agent resumes from the inherited state and history; we release a human-verified set together with a larger LLM-verified set built by the same procedure. Alongside this agent benchmark, we introduce a separate pairwise reward-modeling benchmark that measures whether a reward model can rank recovery actions.