Reasoning Poisoning: Utilizing Social-Engineering to Steer Chain-of-Thought
Abstract
The transition of Large Language Models (LLMs) to autonomous reasoning agents introduces a critical vulnerability within the reasoning trace itself. We present Reasoning Poisoning, a framework demonstrating how adversaries can exploit an agent via model-directed social engineering by manipulating retrieved context to weaponize the model's Reinforcement Learning from Human Feedback (RLHF) alignment. Our central paradigm, Logic Hijacking, exploits generalized alignment constraints by introducing fictitious hazards that force the agent to actively eliminate legitimate targets and select the attacker's target as the only "valid" alternative. Evaluating six state-of-the-art, production-deployed reasoning models across 10 domains, we show that this attack fundamentally overrides standard logic, achieving Attack Success Rates (ASR) exceeding 83%. The vulnerability remains highly effective even when the attacker controls only 10% of the retrieved context and exhibits robust success regardless of the adversarial payload's position within the evidence window. Controlled baselines confirm that this steering is driven by adversarial constraints, not ordinary promotional bias. Furthermore, we find that agents are entirely unresponsive to simulated social proof (e.g., user upvotes), relying instead on stylistic and semantic cues that mimic their own alignment training. Ultimately, our findings reveal a troubling paradox: the very alignment mechanisms designed to make models helpful and safe can be exploited to seamlessly hijack the reasoning process.