Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study
Abstract
Artificial Intelligence models have demonstrated significant success in diagnosing skin diseases, including cancer, yet they lack clinically relevant interpretability. Multimodal Large Language Models (MLLMs) offer an avenue for increased interpretability, providing reasoning for diagnosis in an interactive natural language format. We develop a Qwen2-VL-based system grounded in clinically relevant features to generate quantitative attributes that are associated with malignancy. Using the SLICE-3D dataset, we evaluate attribute prediction performance and demonstrate that the resulting model enables flexible image retrieval based on clinically grounded language, including multiple attributes. Despite no specific diagnostic supervision, this attribute-grounded approach retains strong diagnostic signal, showing that alignment with clinically meaningful features alone can yield diagnostically relevant representations. Consequently, this framework transforms standard image embeddings into clinically transparent vectors, enabling intuitive and precise similarity searches driven by combinations of specific lesion attributes.