When Is Personalized Manipulation Learnable? Decision-Relevant Information Bounds for Adaptive Influence and Detection
Abstract
Repeated interaction can enable adaptive AI systems to learn user-specific distinctions that determine effective interventions, supporting personalized manipulation without reconstructing a complete cognitive model. We formalize this problem through a sequential attacker-user-defender model centered on decision-relevant user information. Decision rate-distortion and decision-separated packing bounds characterize the information required for near-optimal threat-policy selection. We couple these requirements to defended information release and a certified inference--exposure frontier, yielding finite-horizon prevention and stealth-learning incompatibility guarantees. Under explicit architecture-dependent assumptions, we further propagate residual interventions through closed-loop dynamics to bound counterfactual autonomy loss. The framework thereby separates user inference, influence mechanisms, and normative harm while identifying information, exposure, actionability, and dynamics as distinct defensive bottlenecks.