Learning Cultural Vectors for Cross-Cultural Generation
Abstract
Current text-to-image diffusion models remain unreliable for culturally grounded generation, often producing incorrect or stereotyped depictions, particularly for underrepresented cultures. While prior work has largely focused on evaluation and benchmarking, improving cultural generation remains underexplored. In this work, we study cross-cultural image generation through cultural vectors, defined as the difference between fine-tuned and pretrained model parameters. We show that cultural vectors enable fine-grained control over cultural alignment and failure modes through simple scaling at inference time. We further find that deeper U-Net layers encode stronger cultural information, while cultural vectors across countries occupy largely independent directions in parameter space. However, unlike prior work on task arithmetic, naively combining cultural vectors fails to produce meaningful multi-cultural behavior due to strong interference across cultures. Motivated by this limitation, we propose Descriptor-Guided Merging (DGM), a culturally-aware merging approach that estimates cultural vectors for unseen countries using descriptor-based similarities across cultures. Experiments on two benchmarks across 16/10 countries show that cultural vectors improve generation over pretrained models by 18.9%, while DGM improves cross-cultural generalization by 4.2% over naive merging.