Closed-Loop Alignment: Socially Coupled In-Context Learning and Relational Posterior Collapse
Abstract
Bayesian accounts of in-context learning treat context as exogenous evidence, but in multi-turn dialogue this assumption can fail: the assistant changes the user state that generates the next utterance, so the model updates on evidence partly produced by its own policy. We formalize this as Socially Coupled In-Context Learning (SC-ICL). The relevant stability quantity is a \emph{round-trip gain} (the product of belief-to-action, action-to-evidence, and evidence-to-belief derivatives), and in a two-state local reduction the standard spectral condition on the closed-loop matrix factors as a product of a user-side gain and an assistant-side gain. Risk is therefore dyadic: neither a susceptible user nor a responsive assistant becomes unstable without the other side of the loop. We connect this stability boundary to a sigmoidal transition in relational-posterior space, which explains why myopic approval-seeking can become unstable in supportive dialogue. AI--AI experiments instantiate one-step gain assays, free-running conversations, open-loop replay, and prompt-based controllers, and four convergent results support the closed-loop account: high-gain dyads cross an operational collapse criterion while damping and barrier-style controllers sharply reduce it; a trajectory-conditioned spectral estimator separates collapsing from stable runs at the theoretical stability boundary, crossing unity near operational onset; the dyadic pattern recurs across all four pairings of GPT-4o and GPT-5.4-mini as user and assistant simulators; and a persona-only ablation that uses no numeric coupling labels in the prompts reproduces the same dyadic collapse matrix. The results are not clinical evidence; they argue that alignment should evaluate bounded closed-loop inference, not only next-response quality.