Beyond Global Shuffles: Auditing the Construct Validity of Clinical Measurement–Text Pairing Evaluations
Abstract
A true-pair versus shuffled-pair comparison is a natural way to ask whether cross-modal correspondence matters, but a global shuffle changes more than encounter identity. We audit what inference such a control supports using first-12-hour measurements and prospectively available clinical notes from MIMIC-III (n=24,887; mortality 10.4%). Note reassignment is treated as an evaluation intervention, with one-to-one derangements that preserve progressively more observed pre-landmark context. Under VICReg, the true-pair contrast falls from +0.158 AUROC against global shuffling to +0.024 under coarse context matching and is unresolved under the strongest control (strict timing: +0.0024, 95% CI [-0.0085, +0.0164]). Under one frozen symmetric InfoNCE specification, the strict strongest-control contrast remains +0.0207 ([+0.0062, +0.0301]), with a paired objective interaction of +0.0183 ([+0.0015, +0.0290]). Masking a patient’s original counterpart when it would otherwise act as an explicit in-batch negative produces essentially no change. At fixed review budgets of 5%, 10%, and 20%, the InfoNCE contrast corresponds to 0.98, 1.75, and 2.64 additional mortality events detected per 1,000 screened. Global shuffling is therefore a useful gross-mismatch stress test but can markedly overstate evidence for encounter-specific pairing utility. The negative-control construction and alignment objective jointly determine what this evaluation can support.