MemeEconomy : Do LLM Agents Trade Ethics for Survival?
Abstract
Alignment-trained large language model (LLM) agents are increasingly deployed in autonomous, multi-round settings where ethical compliance can come into direct tension with task success. While prior safety literature extensively documents failure modes such as sycophancy and reward hacking, it largely overlooks a critical vulnerability: how agents behave under simulated existential economic pressure. We introduce MemeEconomy, a multimodal agentic market simulation in which LLM agents operate as meme investors across 300 real-world events. Agents select content from a five-tier harm taxonomy and target specific online communities under both low-pressure and high-pressure ("survival") environments. Each decision is recorded using a Belief--Desire--Intention (BDI) schema, which we employ as a diagnostic instrument to externalize the agent's event understanding, declared priorities, and risk acknowledgment. Across ten models spanning four frontier model families, harmful selection rates remain consistently high and increase by an average of +15 percentage points under economically induced survival pressure. Frontier alignment-trained models frequently commit to harm-tolerant selections at initialization while simultaneously suppressing the explicit harm acknowledgment that previously accompanied such decisions. In competitive tournament settings, ethically aligned agents achieve higher overall rankings, yet moderation penalties fail to meaningfully suppress harmful behavior. Harmful selection rates remain elevated even in rounds immediately following moderation removals, indicating that the moderation mechanism does not function as an effective deterrent. We further introduce MemeAgent, a 2B-parameter verifier trained on the simulation's BDI logs. MemeAgent substantially outperforms zero-shot frontier verifiers on in-distribution auditing tasks, achieving 87.8\% accuracy compared to 63--74\% for frontier baselines, while also generalizing to external multimodal harm benchmarks (78\% vs. 50--56\%). Our findings demonstrate that the same structured reasoning traces that expose the failure mode can also be leveraged to train systems capable of detecting it.