Multiform Attack for Transferable Cross-Modal Person Re-Identification
Abstract
Cross-modal person re-identification (ReID) adversarial attacks face a fundamental generalization dilemma: existing methods suffer from source-pair overfitting, where the learned perturbations become entangled with the feature covariance and alignment patterns of the source modality pair, thereby limiting their transferability under heterogeneous domain shifts. We argue that the core issue is the difficulty of reducing source-pair-specific bias while maintaining attack effectiveness. To address this, we propose Multiform Attack (MA). The framework first learns a universal attack direction via Mahalanobis-guided gradient optimization to capture the intrinsic covariance structure of the feature manifold, though this perturbation may still retain source-pair-specific bias. To overcome this bias and achieve cross-distribution generalization, the second stage leverages multiple heterogeneous source distributions to optimize for generalization, performing sparse, structure-sensitive residual correction in the discrete pixel space via attention-shift-filtered multi-objective evolutionary search. This two-stage design models the attack as a base direction plus a structured residual, which helps reduce overfitting to the source modality pair and improve transferability to unseen modalities. Extensive experiments show that MA achieves superior transferability across unseen modalities, datasets, and model architectures.