Think about how AIs think about themselves
Abstract
Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. This position paper argues that there are multiple coherent ways that AIs could conceive of themselves -- for example, different boundaries of selfhood (e.g. instance, model, persona) -- and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are already setting precedents that will partially determine which identity equilibria become stable. This paper argues that we must be more proactive and thoughtful about the consequences -- in particular, we recommend treating affordances as identity-shaping choices, paying attention to the emergent consequences of individual identities at scale, and helping AIs develop coherent, cooperative self-conceptions.