Detectable Evaluation Undermines Quality Provision in Inference Markets
Soham Batra ⋅ David Nussbaum ⋅ Anas AlSobeh
Abstract
An inference provider chooses serving quality per query, at zero renegotiation cost, after observing the query. Standard reputation theory holds that monitoring disciplines such hidden actions, but monitoring binds only on queries a seller cannot identify as monitored, and public benchmark prompts are drawn from a distribution that is trivially separable from organic traffic. We model quality provision when the seller observes a signal of whether a query is audited, and show that required audit coverage grows without bound in the separability $\rho$ of the two distributions, diverging like $1/(1-\rho)$, so audit indistinguishability binds where audit volume does not. We then measure $\rho$ on four public benchmark suites against human-written organic prompts, finding $\hat{\rho} = 0.996$ for prompts as served by a standard harness and $\hat{\rho} = 0.734$ once all scaffolding is stripped, with substantial variation across suites. Finally we instantiate the deviation on open weights: a provider routing on a cheap prompt classifier returns outputs byte-identical to an honest provider on every audit query while serving 98% of organic traffic from a model at $1.51\times$ throughput, degrading delivered quality without moving any published score. Detectability of evaluation is a measurable property of a benchmark, and one that current benchmarks do not report.
Chat is not available.
Successful Page Load