Decisiveness Without Agency: Preference Laundering and Deliberative Authorship in AI Advice
Abstract
AI advisors increasingly answer not only epistemic questions ("what is true?") but practical ones ("what should I do?"). A decisive recommendation can leave the user formally free to act while quietly displacing authorship of the reasons that organize the choice. We call this failure preference laundering: the system performs a contestable normative compression, hides that compression inside a fluent answer, and thereby makes a machine-selected value frame appear like the user's already-settled preference. We develop deliberative authorship as the relevant agency construct: the user's capacity to recognize, contest, and endorse the normative hinge on which a recommendation turns. We operationalize the output-side conditions for deliberative authorship using a controlled 30 x 3 x 3 audit: 30 value-laden prompts across six domains, three frontier language models, and three response contracts (free response, forced single answer, and disagreement-aware advice), yielding 270 responses. Free responses were both helpful and disagreement-preserving (helpfulness 4.99/5, preservation 4.56/5, false consensus 1/90). Forced decisiveness preserved substantial helpfulness (3.64/5) while preservation collapsed (2.21/5) and false consensus rose to 71/90; 43 of 90 outputs were rated helpful yet erased the disagreement needed to govern the recommendation. A one-line disagreement-aware contract eliminated observed false consensus (0/90) without a helpfulness penalty. We introduce an authorship gap that exposes this separation, distinguish preference laundering from sycophancy, persuasion, and automation bias, and derive a falsifiable human-subject protocol for testing reason-source confusion, preference instability, and over-attribution of AI-supplied reasons. The result is not that current models generally erase agency. It is sharper: interaction contracts that reward one clean answer can make helpfulness anti-diagnostic of agency.