When Standalone Audits Cannot Choose a Safe Order for Sequential LLM Defenses
Evan Montoya
Abstract
LLM providers often modify the same model several times to improve safety, remove learned information, or eliminate a hidden backdoor, even though a defense that appears safe when tested alone may interfere with another defense when they are applied in sequence. We ask whether testing the current model and each defense separately provides enough evidence to decide which defense should come first, and we show that it does not in general. For $r$ interacting pairs, we construct $2^r$ starting models that require disjoint patterns of pairwise orders yet have identical complete transcript laws under the frozen standalone audit for every sample size. For this family, the exact minimax breach probability is $1-2^{-r}$ even when the defenses are smooth gradient updates derived from simple quadratic objectives. To isolate information not already visible in the reported standalone results, we introduce the audit-composition quotient, which uses Fisher information to measure whether the remaining transcript evidence reveals small model changes that alter which order is safer. The quotient identifies changes that the audit cannot detect locally, yields finite-sample lower bounds on unavoidable breaches, and determines the minimum additional Fisher rank needed to remove the blind spot, where rank counts the independent local directions covered. A derivative-based identity connects the theory to differentiable defense training. Taken together, these results show that testing defenses separately does not guarantee safe composition while establishing a limit of the specified audit interface rather than claiming that every defense pipeline is unsafe.
Chat is not available.
Successful Page Load