AI Models Can Provably Hide Arbitrary Capabilities
Abstract
Capability evaluations aim to surface dangerous model behaviors before deployment, but their reliability depends on the assumption that hidden capabilities can be elicited. We show this assumption does not hold under adversarial conditions by constructing backdoor attacks that embed encrypted neural networks within host model weights, decrypting and executing them only upon receiving a secret trigger. Using digital lockers, these hidden circuits remain provably hard to elicit and interpret under standard cryptographic hardness assumptions, even given full access to model weights, extending prior work from hiding fixed strings to arbitrary computations. We implement our constructions in PyTorch and empirically validate resistance to supervised fine-tuning on a small chemical reaction prediction task - a proxy for CBRN-relevant capabilities. Our results reveal a limit of capability audits. We release our models to support the development of stronger defenses.