IntentLens: Grounding Underspecified Multimodal Queries for Recommendation via Tool-Augmented Reasoning
Abstract
As intelligent assistants are increasingly deployed in real-world environments, recommendations move from passive preference matching to grounding user intent from underspecified multimodal evidence. In these settings, users often provide only a reference image and a fuzzy natural-language query, leaving crucial visual attributes, personalized preferences, and situational constraints implicit. Existing recommenders either rely on historical interactions enriched with multimodal item features or presuppose fully specified textual requests, which are insufficient to resolve such queries in a single-turn setting. We present \textbf{IntentLens}, a tool-augmented recommendation framework that grounds underspecified multimodal queries into explicit, query-relevant evidence before ranking. IntentLens adopts a \emph{ground-then-rank} design: a shared multimodal LLM orchestrates visual, user-memory, and item-attribute tools to (i) recover latent user intent from the image and language query, and (ii) enrich each candidate with fine-grained, query-conditioned evidence; a lightweight ranker then scores candidates over the grounded representations. To support this new setting, we further construct two benchmarks with underspecified multimodal queries, \textsc{GoogleReview-MMTool} and \textsc{Yelp-MMTool}. Comprehensive experiments show that IntentLens outperforms strong MLLM and retrieval baselines.