SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
Abstract
AI agents are increasingly used to diagnose and mitigate failures of production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to oversimplistic SRE tasks and are unfortunately hard to extend due to bespoke designs. We present SREGym, a high-fidelity, interactive benchmark for SRE agents. SREGym exposes a live system environment built atop real-world cloud-native system stacks, where high-fidelity failure scenarios are emulated through fault injectors. SREGym models the complexity of production environments by simulating (1) a wide range of faults across the stack, (2) various ambient noises, and (3) diverse failure modes such as metastable failures and concurrent failures. SREGym is architected as a modular, extensible framework that orchestrates fault injectors and event emulators across stacks; this framework is key to curating high-fidelity, challenging SRE problems. SREGym shows that the effectiveness of frontier models and agents varies significantly (40+%) in mitigating different kinds of problems. SREGym is actively maintained as an open-source project and has been used by many researchers and practitioners.