SwitchLingua V2: Agent-Driven Code-Switching via Digital Clones
Abstract
Code-switching (CS) involves shifting between two or more languages during a conversation or utterance, a phenomenon often driven by social context and the identity of the speaker. The scarcity of authentic CS data remains a critical bottleneck in developing robust Automatic Speech Recognition (ASR) systems for bilingual communities. While Large Language Models (LLMs) have been employed to synthesize CS data, direct generation approaches typically produce utterances that are pragmatically unnatural and fail to capture the nuanced stylistic variations of real-world bilingual speakers. In this paper, we propose an archetype-driven code-switching dialogue generation framework via agent-based digital cloning. Moving beyond static prompting, our method facilitates autonomous conversations between multi-agent digital clones imbued with specific bilingual archetypes and established linguistic theories. The generation pipeline integrates constrained persona sampling, real-world topic injection, turn-by-turn accommodation dynamics, and multi-dimensional language checking. Applying this framework, we synthesize highly idiomatic bilingual dialogues across 14 diverse language pairs with 260K textual dialogues (1.04M turns of text samples) and convert them into over 1,558 hours of high-fidelity audio data, yielding the SwitchLingua V2 dataset. To rigorously validate our approach and the generated data, we employ a comprehensive evaluation protocol including both advanced LLM assessments and human bilingual expert evaluations. Extensive experiments demonstrate that SwitchLingua V2 produces CS dialogues that are structurally diverse, pragmatically motivated, and socially situated, holding the potential to offer a highly scalable and reliable pathway for advancing multilingual ASR research.