Basis Artifacts in Solution-Level Verification of LLM-Generated Linear Programs
Abstract
Verifiers that read a solver's output to check a generated artifact are usually scored by comparing that output against a reference. This is unsound whenever the reference is not the solver's only correct answer. For linear programs generated from natural-language decision problems, the optimal value is unique but the optimal solution and its shadow prices often are not. A reported discrepancy can therefore record an arbitrary choice among equally valid answers rather than a defect the models force. This paper gives an infeasibility certificate that decides, for a single verifier report, which of the two it is. The certificate quantifies over the whole optimal set rather than over one returned solution. Applied to two verifiers on NL4OPT, it withdraws 96 of a reference solve's 98 unique detections and 31 of 86 for a pair of duality identities read off the same solve. The confound is not confined to small instances. On synthesized programs, in two structurally independent instance families, the share of answer-neutral detections that are basis artifacts rises monotonically with problem size, reaching 93 to 95 percent at four constraints per variable. Every report that survives certification is exactly one structural type: an edit that exchanges two numeric values, leaving the optimal solution unchanged while displacing every shadow price. This error is not confined to constructed inputs either. A production autoformulator commits it in 32.1\% of its erroneous outputs, spanning every slot the edit can occupy. Under an imperfect text extraction, the duality identities degrade less than a syntactic coefficient comparison at every corruption level tested. This reverses the ordering that holds under exact input. Correctness of a verifier's report cannot be read off one returned solution. This paper treats that as a problem of verifier validation and gives a certificate that checks the alternative.