Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
Zeyu Zhang ⋅ Xiangxiang Dai ⋅ Ziyi Han ⋅ Xutong Liu ⋅ John C. S. Lui
Abstract
Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governance that steers behavior without costly retraining. To address this, we introduce the Consensus Clustering LinUCB Bandit (CCLUB), a unified framework for adaptive social alignment via system-prompt routing. CCLUB employs a conservative consensus clustering mechanism: it pools data only within the intersection of utility and safety similarity graphs, effectively preventing unsafe generalization across semantically proximal but risk-divergent contexts. Our theoretical analysis gives an expected regret bound $O\left(d\log T/(p_{\min}\gamma^2\lambda_x) + d\sqrt{MT}\log T\right)$, separating the logarithmic cost of identifying safe consensus clusters from the dominant cluster-level exploitation term. Experiments show that CCLUB improves the safety--utility trade-off and offline deployment gap, and improves cumulative reward by 10.98\% over the strongest non-prototype baseline.
Chat is not available.
Successful Page Load