When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention
Sidi Chang ⋅ Peiying Zhu ⋅ Yuxiao Chen
Abstract
Runtime traces are often treated as transparent evidence about an agent, but a closed-loop policy determines which states are visited and therefore which failures can become visible. We study a simulated hotel-pricing agent whose policy maps time, inventory, and market state to discrete price actions under varying demand regimes. A fault may leave no aggregate trace when the policy rarely visits its affected cells. We treat entry into aggregate-only fault interpretation as a diagnosability decision that precedes scoring or localization. A reference-map gate requires repeated clean-policy support; a matched runtime gate then requires joint support in clean and current streams. Stable signal analysis occurs only after both gates pass. We calibrate false admission on a disjoint clean stream at the physical-component level and model detection by affected clean traffic rather than nominal cell coverage. In a frozen one-shot heldout, 55/72 (76.4\%) regime-component units were reference-admitted, representing 20 physical components; 54/55 then passed matched runtime admission, and the rejected unit abstained. Stable false admission was 0/20, with a one-sided exact 95\% upper bound of 0.1391, meeting the frozen 0.20 criterion. Across 540 repeated unit-arm rows nested in those 20 clusters, affected clean traffic reduced negative log likelihood by 29.3\% relative to cell coverage, a gain of 0.1264 nats per row (cluster-bootstrap 95\% interval $[0.0593, 0.1918]$). Adding mask family and its interaction improved log loss by only 0.0015 nats per row, with a one-sided upper bound of 0.0066, below the frozen 0.01 practical-sufficiency margin. A development audit also found that exact minimum hitting set and greedy selection chose identical supports in 12/12 scenarios because singleton evidence had already resolved the conflicts. The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.
Chat is not available.
Successful Page Load