When Guardrails Look Effective: Construct-Valid Evaluation for Multi-Agent Enterprise Commerce
Peiying Zhu ⋅ Sidi Chang
Abstract
Enterprise agent benchmarks increasingly simulate buyers, sellers, and platform rules to inform deployment decisions. Their outputs can appear decision-ready even when the simulation fails to instantiate the economic roles or isolate the interventions named in the claim. We audit this risk in a multi-turn LLM buyer--seller benchmark for configurable hotel transactions. An initial implementation reported welfare gains of $+87.4$, $+35.0$, and $+28.8$ from two marketplace guardrails across Qwen2.5 models ranging from 1.5B to 14B parameters, but it also changed the offer schema and buyer decision procedure. Holding these components fixed changes the corresponding contrasts to $+7.2$, $-13.9$, and $+23.8$. In a post-hoc repetition of the four largest 14B effects, their mean falls from $+229$ to $+37.6$ with a 95% confidence interval of $[-34.2,109.3]$, while a stronger profit-seeking prompt fails to increase seller profit monotonically. Scripted controls further show that the same buyer-surplus gain can accompany either positive or negative welfare. We introduce a four-gate evaluation contract covering incentive validity, protocol isolation, stochastic stability, and complete welfare accounting. Applied to this case, the original policy effect is invalid under protocol isolation, while the controlled result remains inconclusive under incentive validity and stochastic stability. The benchmark therefore does not support a causal platform-policy recommendation.
Chat is not available.
Successful Page Load