EnvOS: Certified Evaluation of Enterprise Agents in Changing Environments
Niharika singh ⋅ Athira Gopal ⋅ Ankush Kadu ⋅ Arpit Jain ⋅ Aswanth Krishnan
Abstract
Enterprise agents increasingly execute workflows across applications whose authoritative state may change during execution. Existing computer-use evaluations largely assume stable state, obscuring whether agents can detect and recover from such changes. We introduce \textbf{EnvOS}, a framework for generating, certifying, evaluating, and diagnosing agents in changing enterprise environments. EnvOS defines application-independent \emph{phenomena} and environment-specific executable \emph{perturbations}, binding task, world, observation channel, checker, and reward into one evaluation artifact. Executable certification verifies state validity, observation consistency, solvability, failure specificity, and checker robustness; trajectory checkpoints distinguish detection from recovery. We evaluate EnvOS across three multi-application environments using Claude Opus~5, with 32 variants and 93 rollouts. Certification identifies five non-discriminative instances, six checker false-negatives, and a cross-episode state leak. Across certified changed variants, the agent succeeds on only 3/23 rollouts for four recurring phenomenon families. In a controlled \emph{Ghost Commit} experiment, the agent detects every change but recovers in 7/7 pre-commitment trials versus 0/3 post-commitment trials ($p=0.008$). These results show that detection and recovery are distinct capabilities that terminal success alone can obscure.
Chat is not available.
Successful Page Load