EnvOS: A Certified Pipeline for Evaluating Interactive Agents Under State-Changing Phenomena
Athira Gopal ⋅ Ankush Kadu ⋅ Niharika singh ⋅ Arpit Jain ⋅ Aswanth Krishnan
Abstract
Computer-use agents (CUAs) are increasingly evaluated in realistic, stateful environments, but existing evaluations largely assume that the environment remains consistent throughout execution. This overlooks failures caused by changes to authoritative state that occur independently of the agent's actions. We introduce \textbf{EnvOS}, a framework for generating, certifying, evaluating, and diagnosing CUA behavior under such changes. EnvOS defines application-independent \emph{phenomena} and their environment-specific executable \emph{perturbations}, binding the task, environment, perturbation, observation channel, checker, and reward into a unified evaluation artifact. An executable certification pipeline verifies state validity, observation consistency, solvability, failure specificity, and checker robustness before evaluation. Trajectory-level invariants and checkpoints further distinguish detection from recovery. We evaluate EnvOS across three multi-application environments using Claude Opus~5, with 32 variants and 93 rollouts. Certification identifies five non-discriminative instances, six checker false-negatives, and a cross-episode state leak. Across certified changed variants, the agent succeeds on only 3/23 rollouts for four recurring phenomenon families. In a controlled \emph{Ghost Commit} experiment, the agent detects every change but recovers in 7/7 pre-commitment trials versus 0/3 post-commitment trials ($p=0.008$). These results show that detection and recovery are distinct capabilities that terminal success alone can obscure.
Chat is not available.
Successful Page Load