Vendster: An Evaluation Arena to Investigate How Forecasting Tool Quality Affects LLM Agent Behaviors in Buyer-Vendor Negotiation
Charin Polpanumas
Abstract
Time-series forecasters are ranked almost entirely by accuracy metrics, but a forecast's decision value, its effect on the downstream decision, need not track accuracy and is rarely measured. This problem compounds as large language model (LLM) agents increasingly consume forecasters as tools, mediating them through stochastic, opaque policies. We present $\textbf{Vendster}$, an evaluation arena measuring how a forecaster's bias and prediction-interval width change the outcome of a buyer--vendor discount negotiation between two LLM agents. In the reference instance $\textbf{Vendster-M5}$, sampling base demand patterns from the M5 dataset, we evaluate two frontier-closed LLM families (Claude Opus 4.6 and GPT 5.6 Sol) and two open ones (DeepSeek 3.2 and Qwen3-235B-A22B). Model-generated forecasts raise sales target attainment by up to $+17$ percentage points ($p<0.001$), yet standard forecast evaluation metrics such as relWQL and MASE \emph{do not} predict negotiation outcomes at a statistically significant level. To isolate the causal structure, we substitute in synthetic forecasts with controlled bias and interval width. Under-prediction bias significantly improves outcomes for agents that react more to the forecasts---dose-response slopes $-0.387$ and $-0.416$ ($p<0.001$). Interval width has no statistically significant effect at zero bias. These findings indicate that the direction of forecast bias and agent forecast-reactivity are more strongly associated with downstream outcomes than forecast accuracy. Vendster is fully customizable and released under Apache-2.0 at https://github.com/ANON/Vendster.
Chat is not available.
Successful Page Load