DynaProto: Dynamic Prototypical Contrast for Temporally Consistent Object-Centric Learning
Abstract
Unsupervised video object-centric learning requires temporally consistent representations for downstream tasks. Existing methods mainly build temporal consistency on the assumption that adjacent frames share the same object composition. However, this local assumption is insufficient for real videos. Some objects remain semantically consistent over much longer temporal spans, while object composition may also change as objects appear or disappear. Fixed-slot recurrent Slot Attention architectures struggle to accommodate such visibility changes, limiting their ability to learn stronger temporal consistency beyond adjacent frames. We propose Dynamic Prototypical Contrast (DynaProto) to address these challenges. DynaProto introduces a dynamic recurrent Slot Attention architecture that determines whether each slot should be activated at each frame, thereby adapting to changing object visibility over longer temporal scopes. In addition, DynaProto introduces a prototypical contrastive objective that provides stronger clip-level temporal consistency supervision. Together, these two components enable DynaProto to learn more stable slot binding over extended temporal ranges. Extensive experiments show that DynaProto learns more temporally consistent slot representations and substantially outperforms prior methods on object discovery and object property prediction on real-world datasets, including an 11.0\% mBO improvement on YouTube-VIS.