RECAST: dynamic REplay for Counterfactual Agentic Simulation from static Trajectories
Abstract
Continually training and evaluating an enterprise AI agent requires running the agent against the historical environments in which it acted, such as network fabrics, telemetry stores, and databases. However, these environments often cannot be re-run: the incident is over and the live system has changed, leaving only a static historical trajectory. Reviving that trajectory as a queryable simulation model raises three challenges. First, counterfactual coverage: static replay answers only the exact requests on record, so off-trajectory “what-if” queries required for training and evaluation cannot be answered. Second, consistency: an LLM prompted to generate missing tool responses may produce values that contradict other observations in the trajectory. Third, cost: a hand-built simulator is expensive to build and is not anchored to the original incident. We present RECAST, which turns static trajectories into dynamic, queryable simulation models. We use RECAST to reconstruct 2,103 simulation models. On a separate held-out set, RECAST correctly generates responses to 96.3% of tool calls, while an LLM mock gets none right. RECAST reconstructs 95.6% of source trajectories into queryable simulators, compared with 11.37% for a hand-built simulator. It also serves each tool call in 5 ms on average using zero LLM tokens. These dynamic replays improve downstream tasks. Experience mined from an agent’s own trajectories, when injected without weight updates, raises pass rate by 14.2 percentage points on a held-out test set. Separately, training in RECAST simulation models improves pass rate from 44.0% to 52.2% on held-out real tasks.