Designed Autonomy: Evaluative Authorship and Self-Model Lock-In
Abstract
AI can leave the final action to the person while shaping the options, reasons, memories, and self-models through which that action becomes attractive. Prior work already distinguishes action from goal autonomy and studies preference change under AI. Our concern is narrower. A persistent personalised system can become better at serving a learned self-model while weakening the person's practical power to contest, revise, or outgrow it. We call this failure mode self-model lock-in. Unlike mispersonalisation, lock-in can arise even when the model remains accurate: an earlier or endorsed representation may continue to structure retrieval, memory, recommendation, and evaluation after the person wishes to move beyond it. Designed Autonomy frames this as an authority problem over a stored evaluator. We derive design and evaluation implications centred on evaluator provenance, revisable self-models, reflective pluralism, and capability-preserving exit.