Personalized LLM Alignment Should Be Counterfactually Verifiable
Abstract
Personalization shifts the target of LLM alignment from population-level rewards to user-conditioned policies. \textbf{This position paper argues that personalized LLMs must be counterfactually verifiable: models should answer, on demand, three causal queries --- attribution, necessity, and sufficiency --- regarding their inferred user representations.} Absent this property, behavioral evaluation cannot reliably distinguish legitimate adaptation from sycophancy, exploitation, or spurious correlation. Because two models can produce similar aggregate outputs while differing in causal structure (for eg., one truthful, one sycophantic; one respecting protected attributes, one silently conditioning on them) the shift to personalization renders standard benchmarks structurally inadequate. We formalize this via a structural causal model (SCM) of personalized alignment and characterize when these counterfactual queries are identifiable despite the latent nature of inferred user states. We propose a black-box estimation strategy based on rewriting prompts twice to cancel off-target perturbations, and introduce selective counterfactual invariance to bridge the gap between personalization and counterfactual fairness. Ultimately, counterfactual verifiability is both a technical prerequisite for evaluation and a normative standard for the responsible deployment of personalized LLMs.