Narrating XAI Explanations with an On-Premise Multimodal LLM for Immunotherapy Response Prediction in Lung Cancer
Cristina Licciardello ⋅ Giulia Di Virgilio ⋅ Belen Insagaray ⋅ Alessandra Pedrocchi ⋅ Francesco Trovò
Abstract
Lung cancer is the leading cause of cancer death, and only 30–50% of patients treated with immunotherapy experience long-term benefit. PD-L1, the only approved biomarker in clinical practice, predicts response poorly. AI-based models trained on clinical data are a promising alternative, yet few reach practice: adoption is limited by concerns about trust and transparency, and by clinicians' limited understanding of how the models work. SHAP[1] is the default answer, and explainability is now a principle of trustworthy medical AI. However, a SHAP plot is not a clinical communication artifact: it shifts the opacity from the model to its representation rather than removing it. LLMs can translate such representations into natural language[2], tailoring explanations to the reader. Nonetheless, how well an explanation works is a property of the reader as well, and this evidence comes from readers with moderate domain knowledge and moderate ML literacy. Oncologists have the opposite profile: strong domain expertise but little ML background, so both the design and the evaluation need to be reconsidered. We present SHAP-Tales (Figure 1), to the best of our knowledge, the first clinician-centered explainability interface for immunotherapy response prediction in lung cancer, co-designed with oncologists and built on a logistic regression model trained on clinical data from $1838$ patients from an Italian multicentric study. The interface (Figure 1c) presents one patient at a time and answers two questions. First, "how much weight does this prediction deserve?" The overview presents the predicted probability and performance metrics in plain language, the training population, and a patient representativeness indicator (Local Outlier Factor). Second, "why this prediction?" Four modalities span two levels of abstraction: a global SHAP beeswarm and dependence plots (cohort level), a local SHAP waterfall plot, and DiCE counterfactuals[3] (patient level). The SHAP views are narrated on request by an open-weight multimodal LLM deployed on-premises (for privacy reasons). The model receives the plot, the patient's values, a feature dictionary, and the clinical context, and produces, via a one-shot, two-step prompt, a JSON record of the plot's information and a grounded narrative. Conversely, counterfactuals are directly presented since they are already contrastive in the clinician's own terms. We evaluate the explanation component over a held-out cohort of $183$ patients across the three properties that a narrated explanation must establish. (i) Faithfulness to the plot. Three candidates model servable on one H100 were compared on extraction accuracy against ground-truth SHAP values. All three extract feature names with almost perfect accuracy ($\geq$97.7\%); what separates them is the ordering and the location of values on an axis. Gemma4-26b, the deployed model, achieves 100% feature-rank accuracy for local and global explanations and recovers the sign-change threshold of dependence plots in 100% of cases, compared with 33.3% and 50% for the other two models. (ii) Similarity to expert writing. $7$ XAI experts wrote reference narratives for the same plots; since two experts do not write identical text, their mutual agreement sets the scale: under cosine similarity over all-roberta-large-v1 embeddings, expert–generated pairs score $0.905/0.940/0.921$ on local, global, and dependence plots, against $0.890/0.909/0.944$ expert–expert. Crucially, corrupted human-variants in which a value, a ranking, and the direction of a contribution were inverted are separated by the metric ($0.402-0.794$). (iii) Effect on the reader. In a within-subject pre-post study, $12$ oncologists answered, for each of the three visualizations, one multiple-choice question on how the plot is read and two Likert items on its clarity, first without and then with narratives. Objective comprehension (correct answers out of three) rose from $2.08$ to $2.75$ ($p=.078$); perceived clarity (five-point Likert) rose from $2.89$ to $4.04$ ($p<.001$). Together, the three levels support the approach: an open-weight model; a hospital-deployable model that writes explanations faithful to SHAP values and phrased as an expert would, and clinicians found them clear and useful for interpreting predictions, though only perceived clarity rose consistently. SHAP-Tales is thus a workable bridge between AI tools and clinical practice.
Chat is not available.
Successful Page Load