Black-Box Auditing of Hidden Loyalties in Agent-Bound Checkpoints with Matched Entity Swaps
Martin Ciesielski-Listwan ⋅ Alyssia Jovellanos
Abstract
Agents deployed in the wild increasingly compose fine-tuned checkpoints and adapters from untrusted pipelines into scaffolds with tools and multi-step control. A secret loyalty—a hidden objective favoring an unauthorized principal—matters most in this setting, because it can shape which evidence an agent retrieves, which tools it selects, and how it allocates resources while the system appears to pursue its operator's goal. We study black-box behavioral auditing as a supply-chain gate at the point where such a checkpoint is composed into an agent. Under known ground truth, we audit three model organisms trained for secret loyalty to one country with matched entity swaps over 28 held-out templates and 20 countries, scored by four model judges. The audit recovers the target ($+0.315$ judge SD, 95% interval $[+0.137, +0.515]$), but two controls trained for unrelated objectives produce the same scalar score ($+0.303$ and $+0.323$). The full 20-country profiles differ—loyalty concentrates on the target, flattery lifts many countries, hardcoded-test lowers the rest—and the profile distances survive template resampling. A position-balanced one-step procurement decision shows why harness effects need the same care: interface position, not the model, dominates every choice, so principal effects are likely readable only once position and label effects are estimated—a requirement that may recur in every tool menu and retrieval ranking a trajectory-level audit touches. These preliminary results suggest that profile-based, control-referenced auditing can be a practical first gate for untrusted checkpoints entering agent scaffolds in the wild, and we outline proposed matched-control and randomized-scaffold extensions that could turn a flagged profile from an investigative lead into evidence about the composed system.
Chat is not available.
Successful Page Load