Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Abstract
Large language models are increasingly deployed in decisions that require culture-dependent moral judgements, yet they answer as if the whole world thinks with a Western mindset. The Moral Machine experiment showed this is wrong at scale: 40 million judgments across 233 countries reveal that moral preferences are systematically structured by culture, and a model that ignores this variation does not merely underperform, but also imposes one society's intuitions on all others. Existing fixes do not scale to global deployment, as fine-tuning needs per-country preference data and GPU budgets, reward-guided decoding needs per-country reward models, and activation steering needs access to model internals that black-box APIs do not expose. In this work, we focus on this realistic inference-time regime, with no weight updates, no training data, and no internal access. The key observation is that within-country demographic disagreement, not consensus, is the steering signal. When culturally grounded personas agree, the base model is already calibrated. But when they disagree, the spread tells us what to fix and how. We propose DISCA (Disagreement-Informed Steering for Cultural Alignment), which instantiates each country as a panel of four World-Values-Survey-grounded persona agents, converts their disagreement into a bounded, loss-averse correction whose magnitude is set by the panel's variance, and shrinks the correction toward zero when the estimate is unreliable. Across 20 countries and 7 open-weight backbones (2B–70B) from five model families, DISCA reduces cultural misalignment on MultiTP by 10–24% on binary moral dilemmas and 2–7% on open-ended scenarios. Furthermore, a smaller 14B backbone with DISCA reaches lower absolute misalignment than a vanilla 70B model.