Agentic Validation and Repair of Network Automation Programs with the Compiler as Judge
Abstract
Workflow Domain Specific Language (DSL) programs bridge declarative intent and executable runtime behavior. In industry it is common for end to end net work automation to span several repositories rather than a single codebase, and correctness then depends not only on the changed DSL program but also on how those repositories fit together. Validation and repair are hard because a changed program’s behavior depends not only on its accepted language specifications but also on how it invokes the runtime tasks it calls and consumes their results in the other repositories, and on whether the versions it relies on are already deployed. Given a high volume of pull requests, reconstructing this context through human review is slow and misses errors. Because a DSL program can be deployed to hundreds of devices for operations such as mitigation and migration, a single workflow DSL failure partway through execution can halt it and cause a major production incident. We present SentryAI, a deployed agentic validation and repair system in which compilation, decides acceptance. Because the engine compiles each DSL program to the generated code spanning thousands of lines, the compiler reports errors against that generated code rather than the source. SentryAI reasons over the generated code including errors, derives a structured error schema from them, and uses it to localize the fault in the source DSL program and repair it, grounding the fix through retrieval augmented generation with vector search across multiple repositories and a historical repair database keyed by compiler exception codes as a fast lookup path, then reasoning over the results to propose the fix. It recompiles until it succeeds or times out. Beyond compile errors, SentryAI also catches runtime issues from the workflow’s perspective that arise at method invocation, such as mismatches in the required input and output counts, their types, the referenced class names, and the variables passed at the call site. Over six week production windows, SentryAI reviewed 720 commits submitted as pull requests and rejected 23.0% of them, blocking defects that would otherwise have been merged had human review missed them. In a separate controlled experiment on 240 errors derived from production, SentryAI reached 88.6% recall and 92.4% precision