Basis Artifacts in Solution-Level Verification of LLM-Generated Linear Programs
Abstract
Verifiers for LLM-generated optimization models are scored by comparing the solution a solver returns for the candidate against the one it returns for a reference. A linear program has one optimal value but often many optimal solutions, so the returned solution is a choice among them, and this paper shows that comparing returned solutions records that choice wherever the two verifiers disagree. A certificate quantified over the whole optimal set rather than over one returned solution decides whether a report follows from the model or from the basis the solver happened to return. On NL4OPT it withdraws 96 of a reference solve's 98 unique detections, against 31 of 86 for duality identities read off the same solve. The confound is not confined to small instances. On synthesized programs the share of answer-neutral detections that are basis artifacts rises monotonically with constraints per variable, reaching 93 to 95 percent at four, in two independent instance families and at both four and six variables. The paper also identifies the error class that motivates the certificate, formulation errors leaving the optimal value and primal solution unchanged while displacing the entire dual optimal set, and certifies 55 of them. Correctness of a verifier's report cannot be read off a single returned solution.