AnyEdit: A Unified Framework for Speech and Singing Voice Editing with Real-World Environmental Consistency
Abstract
We present AnyEdit, a unified framework for speech and singing voice editing in real-world acoustic environments. Recent unified generative models already achieve strong speech and singing generation in clean conditions, but real-world vocal editing remains challenging because the model must revise content while simultaneously preserving speaker identity, prosody, and the surrounding acoustic scene. A straightforward solution is to directly model noisy edited audio end to end, yet this entangles semantic editing with environmental acoustics and can weaken the core vocal generation capability. We therefore adopt a two-stage design that preserves clean-domain editing ability while deferring acoustic rendering to a separate stage. First, an instruction-guided autoregressive editor predicts clean prosody and content-style tokens. To make this stage robust to degraded inputs, we introduce a teacher-student noise-robust prosody tokenizer that distills clean prosodic tokens from noisy recordings. Second, a flow-matching acoustic model performs in-context acoustic rendering conditioned on the source acoustic context, enabling the edited segment to inherit surrounding noise, reverberation, and background sound while maintaining speaker timbre consistently. To further stabilize the unedited regions, we introduce a unmodified-area-aware supervision mechanism. After pre-training, we also apply direct preference optimization (DPO) using real singing editing pairs mined via dynamic time warping to achieve better singing editing quality. To enable realistic evaluation, we introduce AnyEditBench, the first benchmark covering both speech and singing voice editing under diverse acoustic conditions. Experiments on Chinese and English speech and singing datasets showed that AnyEdit consistently outperforms strong baselines in content accuracy, subjective quality, and background consistency under noisy and reverberant conditions. Demo audios can be found at https://any-edit.github.io.