"Agonistic AI'': Challenging Sycophantic AI as Technology of the Self
Abstract
AI systems participate increasingly in the practices through which people make sense of themselves. Yet current design and evaluation standards incorporate a limited range of philosophical and theoretical work with which to define the scope of possible effects on human selfhood and identity. This is especially relevant for AI sycophancy, a disposition induced by model post-training with harmful social and psychological effects on users. In this paper, we address the gap between philosophical intervention and model development by theorizing sycophancy not as a property of models but instead as a process of self-formation that supplants the relational self. We propose the sociotechnical construct of the interpretive self to counter the “default” sycophantic self, based in descriptive and normative theories that foreground the self as dynamically and relationally constituted. We draw from relevant philosophical, psychological, and technical literatures to offer three design and evaluation targets for the ML community to engage the interpretive self, and we close with next steps towards operationalizing these targets in model post-training. In doing so we demonstrate how theoretical work on human selfhood can inform concrete design and training objectives for ML research.