AI Alignment Can Build Moral Autonomy
Abstract
This position paper argues that moral autonomy is an important and tractable goal in value learning but which current mainstream methods neglect. These methods for model alignment largely improve behavior by supplying models with human preferences, rules, or constitutions. Yet across these paradigms, the authority of the governing principles remains external to the agent. We argue that this external grounding limits the kind of agency current methods can produce. We propose moral autonomy as an alternative alignment criterion, inspired by Kant's account of autonomy. Moral autonomy consists of three necessary capacities: reflective endorsement of one's principles of action, articulation of those principles so they can be enacted and tested, and practical wisdom developed through real-world interaction for sustaining and revising self-legislated principles. Viewed as value learning, these capacities yield a five-level ladder of training paradigms with differing levels of autonomy, from preference-data scaling to real-world interaction. We empirically examine the three capacities, through agent experiments testing whether reflective endorsement improves performance, a survey of 86 frontier LLMs characterizing different levels of articulated normative specification, and multi-agent simulations probing practical wisdom by letting agents develop and revise their own constitutions across changing environments. We close by discussing limitations and alternatives to moral autonomy as an alignment criterion.