When Should a Generative Agent Request Inspection? Cost-Sensitive Information Actions for Bridge Maintenance
Abstract
Generative agents deployed in operational settings must decide not only what answer to produce, but whether the available evidence is sufficient to answer at all. We formulate bridge-maintenance assistance as a cost-sensitive choice between a definitive answer and an information action: targeted inspection or engineer review. The formulation combines task utility, the cost of acquiring missing evidence, and a penalty for unsupported definitive answers. We evaluate this policy on BridgeAgent-Eval, a 5,000-task benchmark derived from U.S. bridge records. On a balanced 1,000-task set, unverified agents achieve task scores of 0.80–0.92 but issue benchmark-policy unsafe answers on 28.7% of trials. Verification removes these answers by routing cases to inspection. A simple Verifier-Routing policy matches AEPV-ReElicit under the declared utility (0.759 versus 0.758) while using 8.43k fewer tokens per trial; a frozen 500-task bridge-disjoint set reproduces the finding. Additional language-model deliberation does not recover structurally missing records, whereas explicit information-action design does.