Semantic Freedom Bottleneck for Domain-Generalized Multimodal Face Anti-Spoofing
Abstract
Multimodal face anti-spoofing (FAS) improves spoof detection by combining complementary RGB, infrared, and depth cues, yet its generalization to unseen domains remains fragile. We revisit multimodal domain-generalized FAS from the perspective of representation freedom. In a controlled CLIP-based post-fusion baseline, increasing visual encoder depth does not monotonically improve cross-domain performance, suggesting that higher-capacity representations may introduce additional degrees of freedom that are not necessarily aligned with real-vs-spoof semantics. To inspect this effect, we introduce the Semantic Dominance Ratio (SDR), a diagnostic statistic that measures the relative proportion of feature variation along the text-defined real-vs-spoof semantic axis versus off-axis residual directions. Motivated by this perspective, we propose Semantic Freedom Bottleneck (SFB), which reparameterizes multimodal features around CLIP text geometry. SFB decomposes each representation into a semantic component along the real-vs-spoof axis and a bounded low-rank residual component for modality-specific evidence. The residual basis is regularized to be orthogonal to the semantic axis, encouraging the residual space to retain complementary modality-specific cues without duplicating the task-semantic direction. This yields a conservative multimodal constraint that preserves discriminative real/spoof semantics while limiting unconstrained non-semantic variation. Experiments on WMCA, CASIA-CeFA, PADISI, and CASIA-SURF under fixed-modality, missing-modality, and limited-source protocols demonstrate consistently strong cross-domain performance.