OS-Omni: A Cross-Platform Benchmark for Generalist Computer-Using Agents
Abstract
Computer-using agents are increasingly expected to complete realistic user goals rather than isolated GUI actions. However, existing evaluations are fragmented across web, mobile, and desktop settings, making it difficult to measure whether the same agent can sustain long-horizon workflows across heterogeneous software platforms while preserving final-state correctness. We introduce OS-Omni, a 1,200-task cross-platform benchmark and execution environment for evaluating computer-using agents across Windows, Ubuntu, macOS, iOS, Android, and Web environments. OS-Omni provides reproducible initialization, platform-specific adapters, trajectory logging, and executable final-state evaluators, and includes controlled service-backed applications for workflows involving account, cart, budget, calendar, message, and other persistent state. Human studies show that OS-Omni tasks require 67.54 interaction steps, substantially exceeding OSWorld references of 15.0 steps. Throughout our experiments, the strongest evaluated agent, GPT-5.5, reaches 47.1% average success, trailing the 88.1% human reference by 41.0 points; on hard workflows, it drops from 77.8% success on easy tasks to 16.4%, while humans remain at 76.3%. These results indicate that current agents remain far from platform-robust, long-horizon computer use, especially when success requires sustained state tracking and durable final-state changes.