Evaluating Agentic AI in Transportation: A Multidimensional Framework and Deterministic Stress Test for Traffic-Simulation Workflows
Abstract
Transportation engineering workflows integrate data preparation, artifact generation, microscopic simulation, calibration, validation, and reporting. Agentic artificial intelligence (AI) may coordinate these interdependent steps, but final-answer correctness alone cannot establish whether the underlying artifacts, tool executions, and engineering decisions are valid, reliable, safe, and reproducible. We introduce a transportation-oriented, end-to-end workflow evaluation protocol that treats the complete request-to-report trace as the unit of analysis. The protocol evaluates task completion, domain validity, tool reliability, robustness, risk control, auditability, and efficiency, and maps the resulting evidence to pass, warning, failure, or human-review states. We instantiate the protocol in a reproducible 40-task deterministic synthetic replay of traffic-simulation workflows. Under fixed and explicitly declared behavior profiles, the configured \RealTwin{} condition with human confirmation attains a 100.0\% task-pass rate, 83.2\% artifact-validity score, 88.2\% synthetic run-status rate, 90.8\% tool score, and 100.0\% critical-action confirmation recall. However, artifact validity falls to 73.7\% on hard tasks even though every hard task passes, revealing that the weighted rubric is too permissive to serve as a deployment gate. A complementary analysis tests the sensitivity of the composite evaluation score to 14 configured backend profiles and yields values of 0.8215--0.9212, but these values reflect declared priors and deterministic variation rather than live provider performance. The framework provides a transparent basis for designing live evaluations with repeated runs, uncertainty estimates, and transportation-domain validators before operational use.