CTE-Bench: Audited Counterfactual Trace Evaluation for Stateful Software Simulators
Abstract
Coding agents built from large language models increasingly inspect and modify stateful services. Before such an agent chooses or explains a patch, it needs a basic engineering skill: predicting how a local code or state change will affect future service behavior. CTE-Bench isolates this simulator-fidelity question from agent policy success. Each item gives the model service source, a factual interaction prefix, a post-prefix source patch or hidden-state overwrite, and a fixed suffix query schedule; an executable oracle supplies the counterfactual response trace. CTE-Bench-Core-v1 contains 255 counterfactual items over six deterministic Python services and 10,200 scored suffix decisions per model configuration. The main score is effect-step value-match (VM): exact response equality restricted to the 2,476 suffix positions where the intervention changes the oracle response. DeepSeek V4-Lite, Kimi K2.5, and Claude Sonnet 4.6 reach 60.7%, 61.5%, and 58.4% effect-step VM under teacher-forced one-step prompting, where each prediction sees earlier oracle suffix responses; hiding those responses while preserving the factual prefix reduces them to 11.4%, 28.9%, and 25.0%. As an additional probe, we run self-conditioned free-rollout on DeepSeek and Kimi; effect-step VM is 27.1% and 33.2%, and exact whole-trace match is at most 1.2%. These probes show that CTE-Bench helps identify dependence on oracle suffix feedback and compounding error in self-conditioned service simulation. We release Core-v1, the executable oracle, metadata, and evaluation scripts, with an evaluation card tying each supported claim to its memory protocol and metric.