Where Does the Low-Cost Model Tier Break in Enterprise Agentic Workflows?
Zhenyu Zhang ⋅ Yuan Ling ⋅ Dingyang Chen ⋅ Xinyang Shen
Abstract
Enterprises running agentic workflows at scale face a recurring cost decision: can a small language model (SLM) replace a state-of-the-art (SOTA) model on repetitive operational tasks? Acting on it requires knowing \emph{where} the SLM degrades, but A/B tests confound model with infrastructure and aggregate scores collapse capability into one number. We present a harness-controlled methodology for building enterprise task environments, instantiated on 1{,}200 paired supply-chain episodes over 103 deterministic tool simulators and scored on 31 sub-metrics (21 rule-based, 10 judged). Comparing Claude Haiku 4.5 to Sonnet 4.5, the aggregate gap ($-0.029$ composite) hides an \textbf{engagement deficit}---calls are constructed comparably ($-2.5$pp argument accuracy, n.s.) while deciding to call one lags sharply: on the $735$ episodes where the harness re-prompts a tool-free answer the SLM still calls nothing in $13.3\%$ of episodes against $1.5\%$ ($95$ discordant vs.\ $8$, McNemar $p{=}5{\times}10^{-20}$), and sequencing trails $7.0$pp---and a \textbf{conservatism--completeness tradeoff}: fewer hallucinations ($3.8\%$ vs.\ $5.3\%$, rule-scored) but less complete answers ($-0.38$ of 5, judged). The deficit is exploitable: escalating only retrieval-required episodes routes $50\%$ of traffic to the SOTA model and recovers $92\%$ of the composite gap, against $56\%$ for rate-matched random escalation. Measuring our own instrument then disciplines our conclusions twice. Re-running one condition moves \emph{absolute} scores by up to $12$pp but the \emph{paired gap} by $0.5$pp $[-2.1,+3.3]$, demoting our scaffold ablation ($+2.2$pp) to suggestive; and all three metrics whose sign reverses against a third model (Opus~4.8) are ones we independently show to be construct-invalid---most sharply calculation correctness, whose scores across four model-runs order exactly by how many numbers each model's own tool outputs contained. Cross-model reversals in agent benchmarks may say more about scorers than about models.
Chat is not available.
Successful Page Load