Beyond Global Alignment: Structured Compositional Reasoning for Vision-Language Models
Abstract
Despite remarkable progress in image-text understanding, vision-language models (VLMs) still struggle with compositional reasoning. In particular, they often fail to distinguish relational direction and attribute-object binding, leading to similar representations for semantically different image-text pairs. This problem mainly stems from the reliance on global image-text alignment, which captures coarse correspondence but overlooks fine-grained compositional structures. Toward this end, we propose an evidence-aware framework, termed Directional Relation Bucketing with Binding Localization (DELTA), to improve compositional understanding in VLMs. The key idea of our DELTA is to improve compositionality of VLMs from two complementary perspectives: direction modeling and binding localization. More specifically, we first leverage a learnable gate to model the cumulative contextual distance between anchor terms, thereby assigning relation concepts to direction-aware discrete buckets. To enrich textual representations with fine-grained relational semantics, we calibrate global semantic representations by incorporating intermediate-layer hidden states. Furthermore, DELTA leverages textual cues to ground visual evidence for object attributes, pulling image patches closer to their matched attribute descriptions while pushing them away from incorrect augmented ones. Extensive experiments on four compositional datasets demonstrate the effectiveness of the proposed DELTA. Our implementation is available at https://anonymous.4open.science/r/DELTA-2D60.