Local Sparsity Enables Unsupervised LLM Safety Detection
Abstract
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Based on this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, and establish a sample complexity bound that scales polynomially in the local sparsity and only \emph{logarithmically} in the SAE's ambient dimension. Finally, we empirically validate our framework on various architectures and datasets, including both capability-testing datasets and safety-specific datasets.