Attention-Derived Local Relational Reasoning for Medical Vision-Language Model Adaptation
Abstract
Generalist medical vision-language models (VLMs) provide broadly transferable image-text representations across medical domains, yet their direct zero-shot use can be limited by the fine-grained visual cues and report semantics that distinguish findings within a specific imaging modality. A practical alternative is modality-level adaptation: adapting a generalist VLM once with representative image-report data from a target modality, followed by zero-shot deployment on unseen datasets from the same modality. We study this setting for chest radiography and introduce ACE-LoRA, a parameter-efficient adaptation strategy that augments low-rank updates with attention-derived local relational reasoning. Instead of relying only on global image-text alignment, ACE-LoRA forms token-centered groups from frozen VLM attention affinities and token similarity, allowing image and report tokens to be refined through group-level context before contrastive alignment. This reasoning layer preserves the frozen backbone as a structural prior while adding fewer than one million trainable parameters. To mitigate false negatives in radiology contrastive learning, we additionally use weak disease-label information to mask semantically overlapping negatives when available. Across CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, ACE-LoRA improves zero-shot classification over specialist and generalist medical VLMs, standard PEFT methods, and full fine-tuning, while component analyses show consistent gains from local relational reasoning in both image and text encoders.