Nash-Sufficient Experiments: Minimax Preference Alignment under Drift and Endogenous Feedback
Manoj Saravanan
Abstract
Preference post-training observes comparisons but deploys a policy. We develop minimax theory for the resulting decision-specific experiment under drift and deployment-selected feedback. An experiment is \emph{Nash-sufficient} when every observational fiber lies inside a fiber of the entropy-regularized Nash map. For $K$ responses these fibers have dimension $(K-1)(K-2)/2$. We derive their fiber-invariant pullback $G_A$, prove $\kernel\cI_A(w)\subseteq\kernel G_A$ as the sharp local observability criterion, and obtain the exact allocation-stable local minimax risk $\frac12\inf_w\tr[G_A\cI_A(w)^\dagger]$. Q-NashTrack adapts to Nash-policy, rather than ambient, variation and is minimax-optimal up to logarithms at the symmetric game. For endogenous comparisons we prove exact sequential characteristic times and a trichotomy; a four-response family has price $1+e^{a/\lambda}$ and infinite price at support loss. A contact-order theorem explains the singular $V^{1/3}T^{2/3}$ phase, and feature-span observability extends the theory to contextual low-rank generalized-bilinear preferences.
Chat is not available.
Successful Page Load