Evaluating and Improving Concept Prediction Models for Histopathological Discovery
Abstract
Histopathology models predict molecular alterations, patient outcomes, and treatment responses with remarkable accuracy. Although interpretability techniques can expose image regions relevant to the model predictions, translating such visual information into human-interpretable biomarkers remains challenging. A promising approach is to use a contrastive vision--language model (VLM) with a pathology-specific concept bank to predict the presence of pre-defined concepts in relevant image patches. To evaluate this approach, we introduce a new concept prediction benchmark and compare three VLMs. Our analysis shows that concept co-activation arising from the geometry of the image embedding space is a major barrier to reliable discovery. We then explore methods to improve concept prediction, including sparsification, ensembling, and finetuning. Our best-performing method substantially improves benchmark performance and demonstrates potential for discovering diagnostic and survival-associated biomarkers.