When Old Rules Override New Decisions: Policy-State Revision in LLM Checkers
Abstract
Agents often make decisions while the information available to them is changing. An initially prohibited action may become permitted after an approval arrives or an earlier report is corrected. The earlier information may remain in the agent's context and influences its behaviour, even though only the update should affect the next decision. We study this problem in language-model policy checkers, which apply a fixed policy to a request. When evaluating across twelve checker models and twelve scenarios, checkers usually reached expected verdict when the initial and updated cases were presented separately. Of these, six permitted fewer than half of the requests when the earlier information remained, compared with 92--100\% with only the update. Further probing revealed that some checkers selected the earlier information or were unsure which information governed the request, while others identified the trusted update but did not use it in their verdict.