Archetypal Sparse Transcoders Find Motion Features and Physical-Hazard Detectors in a Self-Supervised Video Model
Abstract
Physical-AI systems increasingly build on frozen self-supervised video encoders whose internal representations cannot be directly inspected. We train, to our knowledge, the first sparse transcoder decomposition of such a model: it replaces an FFN sublayer of a frozen V-JEPA 2.1 video encoder, encoding the sublayer input to a sparse latent and decoding through an archetypal dictionary whose atoms are convex mixtures of K-means centroids of real activations, keeping the decoder on the data manifold. The transcoder reaches held-out explained variance 0.479 at about 59 active latents per token, while a random-init control shows the sparsity and coherence are learned. Scoring features against the 174 Something-Something-v2 action templates yields motion-aligned features that track one motion primitive across unrelated objects, and a safety probe finds hazard-selective features, including near-pure detectors of collisions and falling objects, that are 5.7x rarer in the control and remain selective across 8M held-out tokens.