When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
Abstract
Interactive simulations are increasingly used to evaluate policies for markets populated by language-model agents. Their outputs can look economic—prices, profits, consumer surplus, and welfare—even when the simulation does not instantiate the economic behavior named in the claim. We audit this risk in a multi-turn buyer–seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B–14B ladder. That implementation also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the same paired contrasts to +7.2, −13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [−34.2, 109.3]), while generation residuals account for 49.9% of the variation in this post-hoc probe. A seller-incentive manipulation check is non-monotone: explicitly increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; guardrails create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract for agent-market evaluation that separates incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returns INVALID or INCONCLUSIVE before allowing a substantive policy claim. Applied to our own case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case study does not establish that guardrails are ineffective; it establishes that their apparent value is unidentified until the simulated agents and protocol pass these checks.