Lost in the Slots: Revisiting Object-Centric Representations in the era of Foundation Models
Abstract
Object-centric representation learning (OCL) has been widely proposed as a principled solution to the binding problem in deep learning, promising improved compositionality, reasoning, and robustness. However, its effectiveness in modern foundation model settings remains unclear. In this work, we systematically evaluate whether explicit object-centric (OC) representations meaningfully improve binding in pretrained foundation models. Departing from prior studies, we assess OC representations in a realistic regime by pairing them with pretrained foundation LLMs and analyze their performance on diverse, open-ended VQA and grounding benchmarks. Our findings challenge the prevailing narrative: slot-based OC representations consistently underperform standard dense patch-based features across multiple benchmarks, including compositional reasoning and hallucination-sensitive tasks. To understand this behavior, we analyze the slot representations through the lens of core binding components, namely scene segregation and object representation,and uncover three key limitations: (i) \textit{inferior scene segregation}, where slot assignments fail to cleanly disentangle objects compared to the implicit grouping already present in pretrained visual backbones; (ii) \textit{inherent information loss} during slot encoding, which degrades downstream performance of VLMs despite adapting them to the resulting feature space of OC; and (iii) \textit{weak attribute encoding}, where slots struggle to preserve fine-grained properties such as category, color, and spatial position. Further analysis reveals that these issues are not solely attributable to slot-based learning: while dense patch-based representations exhibit stronger binding, they too are imperfect. This suggests that explicit object-centric modeling, as currently instantiated, introduces bottlenecks that discard useful information without sufficiently improving structural reasoning. Together, our results expose a fundamental disconnect between the theoretical appeal of OCL and its practical utility in the era of foundation models. They point toward a need for rethinking binding, not as explicit object-factorization alone, but as a balance between structure and information preservation, potentially through new mechanisms that retain the richness of dense representations while enabling more reliable compositional reasoning.