Barriers to Scaling AI Companion Evaluations
Abstract
There is an urgent need for independent evaluations of AI companion systems. Yet, many popular companion products lack public APIs, and little work has studied how researchers can evaluate their user-facing interfaces directly. In this work, we study the prerequisite technical and legal infrastructure required to conduct such evaluations. First, we present four case studies in which we wrote the code needed to directly interact with and test widely used consumer-facing chat interfaces. Second, we survey 12 AI companion systems to assess how common different barriers to independent evaluation are across the broader ecosystem. We find that tens of millions of users use AI companion apps that are functionally exempt from independent oversight.