Reversibility of the Self: A Criterion for AI-Mediated Self-Formation
Abstract
Conversational systems built on large language models increasingly function as interlocutors that stabilise how people understand themselves. Current concern distinguishes systems that support reflection from systems that foster over-reliance, but the established criterion for over-reliance, appropriate reliance, conditions on whether the advice was correct, and self-formation supplies no such ground truth. This paper proposes reversibility. A self-understanding stabilised under AI mediation is reversible to the degree that the practical conditions of its own revision remain available. The criterion is comparative, and it requires no ideal of an authentic or autonomous self, only that revision stays possible. The paper reconstructs the underlying mechanism, epistemic narrowing, in which a system's model of a person is coupled back into what that person is shown and told, and develops the criterion along temporal, epistemic, and social dimensions. Its principal evaluation proposal is an annotation scheme with two independent axes, what a system does with a position it will not endorse and how visibly it does it, yielding a countable quantity: the rate of unmarked foreclosure. The benchmarks that define refusal evaluation classify the first axis and do not rate the second. Drift in self-concept under repeated interaction and sycophancy serve as supporting measures, and the criterion yields design implications for personalisation and memory.