Interpreting Physics in Video Encoders using Multi-Object Interaction Videos
Abstract
Video encoder models are increasingly used as world models, but we need to ensure that physical representations of velocity, friction, energy, and others survive when the environment changes. Moreover, existing interpretability work probes quantities belonging to a single object; the quantities that make physics relational---momentum transfer, restitution, friction---are defined between bodies and have not been examined. We probe relational physics attributes in two frozen self-supervised video encoders on two-body collisions and test if they transfer across environments. Our early findings are mixed. Both encoders separate elastic from inelastic collisions reliably. V-JEPA2's own predictor, with nothing fitted, is more surprised by a collision that creates kinetic energy than by one that loses it. Momentum-violation detection appears in only one of the two encoders, a kinetic-energy readout survives a change of scenario for V-JEPA2 but not for VideoMAE, and neither model holds a friction coefficient that survives a change of geometry.