User Embeddings are Superpositions of Interpretable Behavioral Modes
Abstract
User embeddings are central components of modern recommendation, personalization, and behavioral modeling systems. Though these embeddings can compactly encode rich behavioral information, they are inherently dense and opaque, limiting interpretability. To address this gap, we show that user embeddings can be understood as sparse superpositions of interpretable behavioral modes. This structure is directly predicted by generative models of user behavior for linear embeddings and extends approximately to transformer representations by the superposition hypothesis. Superimposed modes can naturally be recovered by sparse dictionary learning methods, and recovered modes can be automatically named and queried in natural language. On synthetic data with known ground-truth modes, sparse autoencoders (SAEs) outperform classical sparse baselines on mode recovery and produce more monosemantic latents, even when the number of modes exceeds the embedding dimension. Applied to both classical (SVD on open e-commerce) and transformer-based (SASRec/BERT4Rec on MovieLens 1M) user embeddings, the recovered SAE dictionaries are near-orthogonal, and their natural-language labels enable zero-shot audience retrieval over embeddings never trained on text. At matched sparsity, SAE decompositions produce more behaviorally distinct features than classical sparse baselines, with retrieval accuracy comparable to both classical sparse methods and text-embedding baselines and more diverse audience characterizations. The framework is post-hoc and model-agnostic, requiring no modification to the upstream embedding system.