BayesPatch: Causal Auditing of Medical Vision-Language Models via Sparse Activation Patching
Abstract
Reliable vision-language models (VLMs) should not only produce correct semantic associations, but should do so through internal representations that genuinely mediate their visual-language behavior. Observational signals such as saliency, attention, or activation magnitude can identify correlated representations without establishing whether intervening on them causally changes model output. We introduce BayesPatch, a sparsity-constrained activation-patching framework for causally auditing internal VLM representations. BayesPatch searches over transformer layer, token region, and patch strength to identify compact interventions that shift a task-relevant prompt margin while penalizing broad patches. On BiomedCLIP with diabetic-retinopathy prompts, the search recovers the final CLS representation as an expected positive control and identifies an early-layer center region as the strongest spatially restricted non-CLS intervention. In a four-condition validation experiment, DR-to-no-DR transfer produces large positive margin shifts and no-DR-to-DR transfer produces negative shifts, while symmetric same-class controls are substantially smaller. Target-aggregated tests confirm that cross-class effects exceed their matched same-class controls for both interventions. A non-causal activation-magnitude heuristic correlates only weakly with the patching objective and has zero overlap with the top-20 configurations, showing that large representation differences alone do not identify the interventions that most strongly mediate prompt behavior. BayesPatch therefore provides an intervention-based complement to observational grounding analyses while separating activation-level causal mediation from stronger claims of semantic or clinical grounding.