Algorithmic Impact Reveals the Hidden Structure of Alignment
Abstract
When an algorithm makes a decision affecting multiple people, we implicitly have a social choice problem: how should differing opinions about the right course of action be reconciled into a single outcome? We show that if we make welfare consequences of alignment a first-order consideration, this problem can be reformulated as linear optimization on a convex impact space, making it amenable to toolkits from welfare economics and mechanism design. This reformulation provides a clearer understanding of how different protocols shape an algorithm's externalities. Our linear optimization framework uncovers a straightforward solution to strategyproof alignment mechanism—random dictatorship that can be represented as a single parameter, which regulates externalities from manipulation. In addition, our framework allows us to express a family of constrained welfare mechanisms that address participation externalities—how an individual's participation can affect others in the population—and maximize social welfare subject to guarantees on individual outcomes. We illustrate these results empirically using real human preferences over kidney donation assignment, charitable food distribution, LLM responses, and trolley problems.