TraceTriage: A Benchmark for Cost-Aware Stop-or-Continue Decisions in Delayed-Outcome Workflows
Abstract
Current agent benchmarks largely evaluate whether a workflow eventually succeeds, but many real workflows require an earlier decision: whether to continue spending resources or stop and reallocate them before the final outcome is known. In model training, simulation campaigns, and chip implementation, success or failure may be visible only after substantial compute and engineering time. Wrong stops discard runs that would have succeeded; wrong continuations waste downstream compute, tool time, and engineering effort. This asymmetric, risk-constrained intervention problem is not captured by benchmark scores that only measure eventual success. We present TraceTriage, a benchmark for evaluating LLM agents and decision models on cost-aware stop-or-continue policies from checkpoint-local process records. TraceTriage is not a final-outcome prediction benchmark: it evaluates deployable interventions using validation-only thresholding under a false-stop cap and remaining-cost utility. We instantiate the protocol with OpenROAD, an open-source electronic-design-automation workflow that turns hardware designs into placed-and-routed chip layouts through staged tools. OpenROAD is a reproducible substrate, not the benchmark's scope: it provides staged observability, heterogeneous logs and reports, measurable remaining runtime, reproducibility, and verifiable outcomes. The released benchmark uses no-leak run-level splits and instantiates this protocol through 10 OpenROAD design-optimization campaigns with a 50,000-trial budget, retaining 31,572 runs and expanding them to 81,041 checkpoint prefixes with fixed train/validation/test splits. Each example contains only evidence available at a workflow checkpoint, while remaining-cost annotations are used only for evaluation. Together, the results make TraceTriage a capability-boundary benchmark for cost-aware workflow intervention: it verifies actionable signal in checkpoint-local records, quantifies transfer sensitivity, and exposes the unresolved cold-start calibration gap in delayed-outcome, risk-constrained settings.