When Repair Fails: A Preliminary In-the-Wild Audit of User Correction in Grok
Abstract
Most evaluations of deployed language models measure either the first answer or the final task outcome. Benchmarks for multi-turn behavior, sycophancy, and correction usually use synthetic conversations rather than real user interactions. To our knowledge, no prior study tracks what happens after a real user corrects a deployed model in-the-wild: whether the model changes the criticized behavior and whether that change lasts. We introduce a framework for studying user corrections over time. It measures whether Grok acknowledges the complaint, identifies the problem, changes its behavior, keeps that change, and how much effort the user must spend correcting it. We apply the framework to public Grok conversations collected across 21 months. A six-episode pilot found the same pattern in every case: Grok acknowledged each complaint but never fully corrected the behavior. The larger analysis did not show the same uniform pattern. We screened 101 conversations and identified 55 repair episodes. In the first model-assisted coding pass, 21% of acknowledged complaints were not fully incorporated, while some repairs succeeded without any acknowledgment. The main weakness was persistence: the corrected behavior returned in 10 of 18 eligible episodes. Five episodes also showed possible epistemic capitulation, where Grok agreed with the user despite weak or insufficient evidence. Repair failure is an alignment problem that first-response benchmarks can miss. Repair episodes make this problem measurable and provide a framework for further studies of deployed agents.