Towards Multi-Human-Value Alignment via Value Localization in LLMs
Abstract
Large Language Models (LLMs) are increasingly deployed in real-world settings where alignment with diverse human values is essential. However, existing alignment methods are often costly, obscure underlying value heterogeneity, and offer limited interpretability. In this work, we move toward multi-human-value alignment through inference-time intervention by investigating how different values are internally represented within LLMs. We propose a probing-based value localization framework that identifies value-sensitive components, i.e., value heads, which mediate model behavior across multiple moral and normative dimensions such as harmlessness, honesty, and helpfulness. Analyses across multiple LLM families reveal that value heads are universally sparse and exhibit structured interactions, including conflicts among certain values. These components demonstrate clear functional specialization: selectively ablating value heads induces substantial and value-specific behavioral changes. Building on these findings, we introduce an inference-time intervention strategy that enables controllable adaptation of LLMs to single or multiple human value systems, with theoretical support for its effectiveness. Experiments on human-value benchmarks demonstrate that our method improves flexible and faithful value alignment while maintaining overall model performance, highlighting mechanistic interpretability as a foundation for socially aware, pluralistic AI.