What Do Multimodal Conflict Detectors in Pathology Actually Detect?
Abstract
Multimodal pathology models fuse a tissue image with its written report, so a mismatched pair can quietly corrupt their predictions. The usual safeguard test swaps reports and reads a high conflict score as evidence that the model catches clinical contradiction. We show this reading is unsafe: a swap can be flagged merely because the new report names a different tumour, follows a different template, or belongs to a different patient. Using a ladder of controlled comparisons over six frozen encoders, we ask what a conflict score actually responds to. A lung benchmark first confirms that swapped reports both harm prediction and are easy to flag, yet a plain count of diagnosis phrases flags them just as well. Holding each image fixed, we then separate a broad different-patient response from the added response to diagnostic class, and compare a parameter-free cosine score with a small learned detector on the same embeddings. The detector sharpens class separation for five of six encoders but leaves the different-patient response uncertain, and its advantage rarely survives graded text replacement or transport to other cancers. What a conflict detector detects therefore depends on the representation, the scoring rule, and how the disagreement is constructed, so a swap score alone cannot certify clinical conflict detection.