Where Should the Code Go? Locating Foundation-Model Errors in Bitemporal Inventory Simulation
Abstract
Operational reports carry two clocks: when an event happened and when its report arrived; later reports can revise earlier ones. A foundation-model component that maintains the state supported by the reports visible at a cutoff, then rolls a supplied scenario forward, can be helped by exact code before it (normalize the evidence), after it (execute from its emitted state) or inside the call (a code tool), or simply asked to reason harder. We ask which remedy removes which error, for which models, on an exactly solvable lost-sales inventory environment: 36,672 independently rescored requests over twelve configurations of closed and open-weight models, every configuration under one tool-calling interface with the forced-answer-call interface as a control, all under protocols frozen before the calls. Weak models fail in separable ways. Canonical input removes report-selection error, exact continuation removes rollout error, neither repairs the other, and which error dominates differs by model; strong models are at ceiling on raw reports. Asking for more reasoning is no substitute and is interface-dependent: under a forced answer call it collapses GPT-5.4 mini and DeepSeek V4 Pro while helping them under automatic choice, it silently disables reasoning for DeepSeek V4 Flash, and at high effort gpt-oss exhausts its output budget. The contribution is a controlled component diagnostic, released with every request, and a warning that the tool-calling interface confounds reasoning evaluation.