Character Mixing for Video Generation
Abstract
Imagine Mr. Bean stepping into Tom and Jerry--can a video model generate interactions that stay faithful to each character's identity, behavior, and style across worlds? We present MiMix, a unified framework for multi-character text-to-video generation that goes beyond visual appearance to capture character-specific personas generalizing to unseen scenes and partners. To evaluate this, we introduce the Cross-Character Generalization Benchmark, a social stress test placing each character in novel scenes with previously uncoexistent partners. MiMix combines Cross-Character Embedding (CCE), which disentangles identity and behavior via text-aligned annotations, with Cross-Character Augmentation (CCA), which synthesizes cross-style training data while preserving native appearance. Experiments show consistent gains in identity fidelity, interaction quality, and style consistency over prior personalized and general video models.Additional results and videos are available on our project page: https://mi-mi-x.github.io.