When Correctness Feedback Erases Evidence: Phase-Separated Completeness Auditing in Verus
Abstract
Formal verification establishes that a program satisfies its written contract, but not that the contract captures what its author intended. We study completeness auditing as a search for a concrete program that passes a weakened contract while violating intended behavior. This objective creates a conflict with familiar correctness feedback: verifier messages and visible tests can steer the model away from the wrong program that would expose the missing obligation. Across 12 Verus/Rust tasks, four model families, and three labeled repeat episodes, no-feedback adversarial discovery certifies 85 of 144 outcomes, while IntentShield, which provides corrective feedback during discovery, certifies 39. The difference favors no-feedback discovery on 11 of 12 tasks and is 31.9 percentage points on average (task-bootstrap 95% interval [21.5, 38.9], p=0.0063). A phase-separated execution of the same no-feedback protocol independently certifies 85 of 144 outcomes before freezing the selected candidate for repair. In a separate 16-task evaluation, phase-revised specifications yield 39 of 48 successful implementations, compared with 30 of 48 for IntentShield-revised specifications and 14 of 48 for the original weakened specifications. The effect did not reproduce in the available Dafny or Lean evaluations, and the downstream study lacks a Direct-Revised baseline. Within the studied Verus setting, the result supports a simple rule: preserve and certify the adversarial witness before beginning correctness-oriented repair.