CrossID: Cross-Supervised Spatio-Temporal Gated Fusion for Personalized Portrait Generation
Abstract
Personalized portrait generation aims to synthesize portraits that align with both the identity of the reference face and the semantics of the text. However, existing methods suffer from two critical challenges. First, they exhibit a trade-off between ID fidelity and text controllability. We identify that this limitation stems from the self-supervised training, where the model learns a trivial mapping rather than understand identity. Second, despite exploring various encoding strategies, recent approaches still tend to discard fine-grained facial details. To address these challenges, we propose CrossID-11M, a large-scale curated portrait dataset comprising 11 million images with high-quality annotations, paired with millions of identity-preserving face video clips. Furthermore, we introduce CrossID, a novel framework integrating a lightweight GatedFaceAdapter and a cross-supervised training paradigm. Additionally, we establish CrossID-Bench alongside new VLM-based evaluation metrics. Extensive experiments demonstrate that CrossID successfully breaks the trade-off between ID fidelity and controllability, achieving state-of-the-art performance by producing photorealistic images with faithfully preserved ID details.