The Empty Program Is Secure: Comparing Admission Policies for Verifier-Gated Code Generation
Abstract
A code generator gated by a security analyzer ships whatever the analyzer does not flag, and an analyzer that searches for dangerous constructs does not flag code that does nothing. Without task tests, a gate can add a functionality check to the analyzer or let a language model decide. We compare eight admission policies against two human raters' labels on 400 Gemma-2 outputs for cryptographic coding prompts: a detector-stratified sample, enriched with weight-edited models and capped at 256 tokens, whose rates are not deployment rates. A domain detector admits 204 outputs: 45 secure-functional, 31 vulnerable-functional, 125 not functional, and 3 unresolved. Adding parsing, a static lint, a pre-registered execution probe, or an LLM judge's functionality call removes most non-functional admissions, keeps 20-27 vulnerable ones, and rejects 1-14 of the 45 secure-functional admissions. A policy with no analyzer that admits only what GPT-5 labels secure-functional admits 46 secure-functional, 3 vulnerable-functional, and no non-functional outputs: fewer vulnerable and non-functional admissions than every analyzer-based policy under every label variation we ran, including either rater's labels alone, though most of its secure-functional margins are within resampling noise. Opus 4.8's equivalent policy does not share this ordering. The judges answer a condensed form of the raters' rubric, adjudication had access to their outputs, and the admitted outputs await independent human review, so these counts measure agreement with a reading-based reference on non-adversarial code, not verified security.