Assurance Boundaries in AI-Generated Software
Abstract
Autonomous coding changes both the rate and the structure of software development. A coding task specifies behavior, a generator produces a candidate, and tests or formal verifiers provide evidence about that candidate. In an autonomous setting, however, the code, requirements, checking properties, and evidence may all change during the same trajectory. A verifier can therefore be correct about the property it checked while the resulting acceptance decision is still wrong for the task. We study this gap by separating four assurance boundaries: the obligations derived from the task, the properties or oracles used to represent them, the evidence produced for those properties, and the admission rule that determines whether a candidate may be committed. Using a builder-independent assurance kernel as an experimental instrument, we evaluate both controlled boundary faults and dynamic changes to assurance state. Across two multi-obligation programming tasks, repairs are boundary-specific: strengthening the wrong boundary can drive verifier success to 100% without improving task-correct acceptance. In dynamic traces, artifact identity alone is also insufficient because task, obligation, or property state can change while the code does not. Together, these experiments show that assurance has its own state structure: task obligations, checking properties, evidence, and admission are distinct coordinates, and improving one does not in general repair a failure in another.