Mixture of Attribute-Aware Attention Experts for Fine-grained E-Commerce Composed Image Retrieval
Yufei Ma ⋅ Zihan Liang ⋅ Zhipeng Qian ⋅ Huangyu Dai ⋅ Lingtao Mao ⋅ Ben Chen ⋅ Chenyi Lei ⋅ Wenwu Ou
Abstract
Composed Image Retrieval (CIR) enables users to express specific search intent by combining a reference image with manipulation text. Despite its promise for e-commerce, existing methods ignore readily available item textual descriptions, suffer from multi-attribute confounding where samples exhibit uncontrolled variations in unrelated attributes, and lack explicit fine-grained attribute modeling. We propose $\textbf{MoA}^{3}\textbf{CIR}$, featuring an asymmetric five-tower architecture that incorporates item descriptions on both query and target sides, a self-reflective data augmentation pipeline generating controlled attribute modifications through LVLM-driven refinement, and a Mixture of Attribute-Aware Attention mechanism with grouped multi-head attention and dynamic routing coupled with a two-phase training strategy. Besides, we contribute ECCIR, the first dual-modal e-commerce CIR benchmark with 110K samples featuring authentic item descriptions. Extensive experiments demonstrate that \method~achieves SoTA performance on ECCIR, FashionIQ, and Shoes datasets, with substantial improvements over existing methods. Code and datasets will be made publicly available.
Chat is not available.
Successful Page Load