Diffusion Loss as an Anomaly Signal for Out-Of-Distribution Inputs
Soham Chowdhury ⋅ Shivam Arora
Abstract
Recent research demonstrates that generative meta-models, specifically Generative Latent Priors (GLPs), can effectively approximate the natural activation manifolds of Large Language Models (LLMs) by learning their residual stream activation distributions. Out of-distribution inputs to LLMs cause their activation to go out of their natural manifold. This paper shows how the Generative Latent Prior's diffusion loss can be used to reliably detect these out-of-distribution inputs to LLMs.
Chat is not available.
Successful Page Load