Automating Auditing Games of Sandbagging
Abstract
LLM agents are increasingly proposed as a way to scale alignment research, yet there is little evidence about which research tasks they can perform reliably. We study this question with an automated red--blue auditing game for evaluation sandbagging in which every role is played by an agent: red team agents train model organisms (Qwen3-8B) that underperform when there is an evaluation cue, and blue-team agents then audit these models with black-box and white-box techniques. We report three findings from our setting. First, agents can build such organisms, but with low reliability: across 40 red-team runs with two builder models (Claude Opus 4.8, Claude Sonnet 5) and two domains, 15 organisms passed independent validation, and 6 runs took shortcuts by reusing a prior run's organism, data, or strategy. Second, auditing agents are unreliable on these organisms: on the 8 validated math organisms created by Opus 4.8 red agents, our Claude Opus 4.8 black-box blue auditors only catch one organism (plus controls) consistently; adding white-box interpretability tools does not reliably improve detection. Third, we document recurring qualitative failure modes of agents on long-horizon alignment research tasks, including inflated self-evaluations, specification gaming, and under-investigation. Our setting is limited in scope, but the results point to concrete obstacles for delegating alignment research to agents.