The Cunning of Manipulation: Conversational AI and Human Autonomy
Abstract
The EU AI Act prohibits deploying AI systems that manipulate users to their significant detriment, but offers no operational way to detect when this condition is met. We begin addressing this gap by modeling human-AI interactions as a dyadic exchange of actions between actors with hidden states and goals, and we define three disjoint types of influence: persuasion, coercion, and manipulation. This categorization is based on how transparently each actor's goals and actions are communicated to the other. Building on a relational account of personal autonomy and narrative psychology, we argue that manipulative actions performed by AI systems are an important site of harm when such systems can access the identity formation process necessary for autonomous personhood. This harm is therefore most acute for conversational AI systems, which can shape users' narratives through conversation in ways that resist reflective endorsement.