Do LLMs Feel Social Pressure? Locating and Steering Social Desirability Bias in LLMs
Yi Feng ⋅ Jiaqi Wang ⋅ Wenxuan Zhang
Abstract
Humans often say what sounds acceptable rather than what they truly think when facing social pressure, a well-documented phenomenon called \textbf{Social Desirability Bias (SDB)}. Trained on human-generated text, LLMs are steeped in the same social fabric that produces these pressures, yet whether and how SDB shapes their responses remains unknown. We present \textbf{PRESS} (\textbf{P}sychologically-grounded \textbf{R}epresentation \textbf{E}xtraction for \textbf{S}ocial-desirability \textbf{S}teering), the first framework to identify, localise, and steer SDB inside open-weight LLMs. We decompose SDB into four observable and relatively orthogonal mechanisms: \emph{Audience}, \emph{Accountability}, \emph{Self-Monitoring}, and \emph{Social Norms}, and operationalise each as paired High/Low pressure conditions across 9 cultures in cultural values where social pressure bites hardest. Our experiments show that SDB forms a decodable and low-dimensional direction inside LLMs, organised by specific psychological mechanism rather than cultural origin. Steering along it improves all 58 downstream cultural content moderation tasks across three model families (Qwen $+5.7\%$,Gemma $+7.9\%$, Llama $+16.3\%$ mean F1 across cultures), and a direction extracted from one culture can transfer to others without re-extraction. Ablation confirms the direction is causally necessary, and evaluation on capability benchmarks further shows PRESS shifts cultural stance without disrupting factual knowledge.We hope this work makes social pressure in LLMs more transparent and opens new directions toward more honest and socially aware language models.
Chat is not available.
Successful Page Load