Reliable in Isolation, Vulnerable in Interaction: Misleading History in LLM Climate-Denial Assessment
Abstract
Large language models (LLMs) play an increasing role in simulating social dynamics, which requires the safety alignment abilities of LLMs to transition seamlessly into interactive environments. One key problem here is that the stimuli from the environment might not always be trustworthy, meaning that the consequent model responses might deviate from expectations, thereby polluting downstream analysis. This study reveals that, even with stronger models, manipulating a simple piece of misleading history can penetrate the safety alignment and cause serious performance degradation in the climate-denial claims detection task, which originates from a well-studied and documented topic. This key finding urges the further development of safe AI for trustworthy interactions.