Towards Native Agent Evaluation: A Minimal Interface for an Agentic OS
Abstract
Agent evaluation today is reimplemented inside every application, in ad-hoc formats that neither transfers between projects nor composed into shared knowledge. We argue it should instead be in the Agentic OS layer as a stable interface of four primitives: Record, Metric, Suite, Gate - specified against four properties an operating system should provide: startable, composable, repeatable, understandable. We present a reference implementation of these primitives that couples evaluation to the agent lifecycle from early development through production. We showcase our learnings from building and using our evaluation system, both internally and in production deployment at a national telecommunications operator. We also share some design principles, intended as a basis for the community to build evaluation system on shared standard practices than private conventions.