Uncovering Hidden Objectives via Transcoder Circuits
Abstract
During training or fine-tuning, large language models (LLMs) can develop unintended goals or exhibit various forms of misalignment. Furthermore, they can be intentionally trained to pursue hidden objectives, harbouring secret loyalties or backdoors that can be introduced by even a small number of poisoned documents in the training corpus. While robust and scalable methods are urgently needed to audit model behaviour, black-box approaches frequently fail to elicit the true underlying goal, particularly when models are trained to conceal it. In this work, we investigate automated methods for auditing models specifically trained to exhibit reward model sycophancy. We demonstrate that transcoder-based attribution enables the robust identification of interpretable differences between the base and fine-tuned models, successfully scaling this analysis across architectures ranging from 4B to 70B parameters. We introduce an automated auditing agent that leverages these internal features to successfully detect the embedded goal, whereas black-box evaluations and activation-difference analysis reveal only superficial behavioural artifacts. Finally, we show that a less capable auditing agent fails to reliably identify the hidden objective even when provided with identical internal evidence, underscoring the necessity of highly capable models for auditing complex, deceptive behaviours.