$\text{PartConcepts}$: A Unified Mechanism for Fine-Grained Part Localization and Generation
Abstract
While text-to-image (T2I) diffusion models exhibit strong semantic disentanglement at the object level, they struggle to localize fine-grained object parts. Although the latent visual features in these models are sufficiently fine-grained, the text–image interaction (cross-attention) fails to effectively exploit this information. This results in coarse and ambiguous localization, particularly for spatially distinct components such as left and right limbs. To address this limitation, we introduce PartConcepts, a mechanism that encodes textual part descriptions into compact, learnable representations. These PartConcept tokens are trained to selectively attend to their corresponding part regions in the image. To evaluate the part-level understanding of PartConcept tokens, we first probe its effectiveness for part segmentation. We then evaluate whether the improved localization enables fine-grained instruction following in T2I generation. On the standard Open-Vocabulary Part Segmentation (OVPS) benchmark, our method outperforms dedicated segmentation baselines, establishing a new state-of-the-art. We further evaluate our method on the more complex Pascal-Part benchmark for part instance segmentation, which contains fine-grained spatial annotations (e.g. left lower arm). Here, we outperform all baselines by a very significant margin. Finally, our qualitative and quantitative results show that this strong localization of PartConcept tokens directly enables fine-grained instruction following in T2I generation: the textual attributes (e.g., color) inherently bind well to corresponding PartConcept tokens, without explicit attribute control mechanisms. This indicates that PartConcept tokens compose seamlessly with other text tokens, highlighting their applicability for fine-grained text-based generative control.