One Fact Apart: Hidden Failures in How AI Assistants Follow Rules
Zahid Syed ⋅ Aidan Lau ⋅ Moulik Jain ⋅ Cooper Lee ⋅ Shivank Garg
Abstract
AI assistants often receive written rules that tell them whether to answer a user's request, ask for more information, refuse, or redirect the user. The required choice can change with a single detail, such as whether permission remains valid, an exception applies, or needed information is missing. We test whether models follow these changes using OneFact, a benchmark of $100$ three-case ``folds.'' Each fold contains three closely matched versions of the same situation. In all three, the assistant receives the same written rules and nearly the same user request. Only one important fact changes, and that change may require the assistant to respond differently. Across five model families and two prompting conditions, item-level agreement, the share of individual cases that match the reference response, ranges from $40.3%$ to $56.0%$. Complete-fold agreement, which requires a match on all three cases, ranges from only $1%$ to $14%$. In $43.9%$ of mismatches, where the model selects another response, it still cites at least half of the rules supporting the reference response. A reasoning prompt, which asks for an explanation, changes $29.2%$ of paired responses and produces substantially longer outputs without consistently improving agreement. These findings show that scoring each case separately can hide whether a model responds consistently when one important fact changes.
Chat is not available.
Successful Page Load