FinAl: Fine-grained Alignment for Detail-Preserving Medical Vision-Language Pretraining
Abstract
We present a vision–language pretraining (VLP) framework for the medical domain that focuses on preserving fine-detail clinical information, such as disease severity, which plays a critical role in downstream medical applications. Despite recent advances in medical VLP, existing approaches often struggle to retain such subtle semantics. This limitation mainly stems from the use of contrastive learning schemes that treat supervision as hard one-hot labels, overlooking the ordinal and overlapping nature of clinical findings and thereby introducing false negatives. To overcome these challenges, we propose FinAl, a fine-grained alignment approach that incorporates soft alignment targets into vision-language pretraining. These targets are derived from structured radiology reports and encode relative semantic relationships among clinical descriptions, allowing fine-grained linguistic information to be transferred into the visual representation space. Extensive experiments show that FinAl improves performance on fine-grained clinical tasks, including severity grading and severity-aware retrieval, with additional evaluations on challenging fine-grained scenarios further supporting its robustness. At the same time, the proposed method maintains competitive performance on standard coarse-grained clinical benchmarks.