Detecting Research Sabotage from Agent Behavior
Abstract
LLM agents can corrupt research by placing deliberate errors in work that otherwise appears useful. We ask whether consistency checks developed for sandbagging detect research sabotage, and whether additional attempts, judgments, or follow-up questions improve detection. An actor follows either honest or attack instructions, and an evaluator distinguishes the conditions from the actor's work. We run live agents on APPS programming and SAE feature-interpretation tasks and evaluate existing transcripts and repositories from three further datasets. We find that Cross-Context Consistency (C\textsuperscript{3}) checks transfer from sandbagging detection to sabotage detection, but the compared output matters: APPS code is more informative than actor explanations, and the checks also distinguish instructed attacks on SAE Interp. Resampling agent trajectories does not always improve sabotage detection: taking the maximum monitor score across up to ten APPS attempts improves AUROC from 0.616 to 0.932 for one actor, whereas aggregating Ctrl-Alt-Deceit's existing agent trajectories gives little gain at the same requested false-positive rate. Rejudging and follow-up questions do not improve AUROC in our tests. Instructed model organisms expose cues that favour evaluators: attack transcripts often state their objective, transcript statistics distinguish assigned conditions, and our monitor detects warned attackers more easily than unwarned attackers. These results favour collecting and comparing new task behaviour before treating agent work as reliable research evidence, while showing why stronger detection claims require more realistic instructed model organisms.