ST-Bridge: Bridging Sketch and Text with Large Language Models for Coarse-to-Fine Image Retrieval
Abstract
Sketch-Text image retrieval aims to find natural images using both free-hand sketches and text descriptions. Although these two modalities provide complementary information, the task is challenging due to their large representational differences and the difficulty of matching fine details in complex scenes. Most existing methods rely on simple feature fusion or global alignment, which often fails to preserve important discriminative cues and leads to inaccurate retrieval results. To tackle these issues, we propose ST-Bridge, a high-precision sketch-text image retrieval framework that enables reliable coarse-to-fine matching through an effective semantic bridging mechanism. Specifically, we introduce LLM as a semantic bridge to generate sketch representations that are closer to the textual semantic space. We further enhance these representations and the original sketch features via gated attention and feature normalization, substantially reducing the sketch--text modality gap. Building upon this, we design a coarse-to-fine alignment strategy that supports accurate sketch--text image retrieval. Extensive experiments on scene-level sketch--text image retrieval benchmarks demonstrate that ST-Bridge significantly outperforms existing methods and achieves state-of-the-art performance, validating the effectiveness of large-model-based semantic bridging for cross-modal retrieval tasks. Our code and models will be released upon acceptance.