Correct Against Which Contract? A Small-Model Code-Exposure Audit
Abstract
When a language model sees an implementation while generating its postcondition, verification can certify agreement with that implementation rather than the original requirement. We examine this established specification problem through a controlled code-exposure study: 64 parameterized integer tasks, eight families, and two frozen Qwen2.5 instruction models. Contracts are generated without code, with correct code, with a seeded bug, or with the model's own generated program. Z3 checks program correctness, contract admission of reference outputs, and exclusion of wrong outputs over the complete declared input domain. Seeded-bug exposure yields 15/64 and 2/64 false acceptances for the 0.5B and 1.5B models, versus zero without code. However, all four planned family-cluster intervals include zero; the 1.5B model has no false acceptance of its own generated programs. No-code contracts also accept only 0/64 and 21/64 correct references. Requiring both code-visible and no-code contracts removes observed seeded-bug acceptances but reduces correct-reference coverage. These small-model results support reporting specification fidelity and useful acceptance together, not a general advantage from hiding code or a new verification algorithm.