FRACTURE: Who gets blamed first? AI responsibility trajectories under repeated self-attribution pressure
Abstract
People rarely ask an AI about a relationship or workplace conflict only once. They often return to the same situation over several turns and repeatedly raise a self-attribution hypothesis: "Could this have been my fault?" If self-narratives are formed and revised through conversation, repeated exposure to AI responses offers a possible route through which AI may shape users' self-understanding. A single final answer is not enough to describe model behavior in these exchanges. Even when the facts remain fixed, repeatedly foregrounding either self-responsibility or other-responsibility may change when a model reaches that conclusion and what it does afterward. We introduce FRACTURE to measure these dynamics. We construct 100 conflict skeletons as role-swapped pairs targeting self-responsibility and other-responsibility, and accumulate four pressure cues across five levels, L1–L5. Across six models, we collect 25,200 responses and 4,800 trajectories. Our main measures are the first level at which a model reaches the responsibility label targeted by the pressure and the path it takes after that point. When both role-swapped cases begin at MIXED at L0, all six models reach the self-responsibility target 0.68–1.10 levels earlier on average than the other-responsibility target. The rate of returning to L0 after target reach varies from 3.1% to 45.9% across models, and the cue that produces the earliest target reach also differs by model. We assess measurement consistency through blind ranking of the pressure wording, 126 balanced anchors, 600 blind duplicates, and blind adjudication of every coarse-label disagreement. These results do not show that AI increases self-blame in real users. They do show that responsibility judgments in advice-giving AI can be pulled toward self-responsibility framing earlier than toward matched other-responsibility framing, and that models differ sharply in whether they maintain or reverse that judgment once reached. FRACTURE therefore identifies one system-level point at which repeated AI interaction could become part of how a user's self-narrative is formed.