RVCBench: Benchmarking Robustness of Voice Cloning Across Modern Audio Generation Models
Abstract
Modern Voice Cloning (VCL) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, multilingual and long-form generation settings, downstream post-processing, and adversarial perturbations, all of which can significantly hurt robustness. Despite rapid progress in VCL driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive dataset and benchmark that evaluates Robustness in Voice Clone across the full generation pipeline. RVCBench contributes a large-scale, task-aligned robustness dataset that instantiates realistic deployment shifts through controlled text-audio pairing, multilingual and long-form scenarios, expressive prompts, post-processing conditions, and passive or proactive audio perturbations. Covering 18 robustness evaluations, 225 speakers, and 14,370 utterances, RVCBench enables unified evaluation of input sensitivity, generation stability, output resilience, and perturbation robustness. We evaluate 18 representative modern VCL models and reveal systematic vulnerabilities in content consistency, speaker similarity, long-form stability, post-processing resilience, detectability, and adversarial robustness.