Alignment and Candor in AI Delegation
Abstract
An AI agent that acts on your behalf has to learn what you want. But if it also weighs the interests of the people across the table, you may hesitate to tell it everything. We study how a delegate's objective changes its own user's willingness to share private information. Applying standard tools from strategic communication and mechanism design, we show that loyalty supports honest disclosure, while broader social concern can undermine it. When the surrounding institution can be redesigned, every equilibrium outcome achievable with socially concerned agents can also be achieved with loyal agents and truthful user reports. Social concern can therefore be placed in the rules while preserving faithful representation. When the rules are fixed, an intermediate degree of social concern can instead be best. These results suggest a change in how we evaluate AI delegates: the information users provide should be understood to be endogenous to the agent's objective which, in turn, is shaped by various alignment treatments.