FlexCover: Flexible Cover Song Generation via Symbolic Lead Sheet Control
Abstract
A cover song re-renders an existing piece by preserving its tonal content—melody and chord progression—while reshaping other musical attributes. This suggests a natural formulation for generative cover synthesis with pretrained music language models: explicitly control tonal content, while specifying lyrics and style through text. Existing approaches condition on frame-level features, tightly coupling generation to the reference’s timing and structure, and limiting flexibility such as tempo variation, structural rearrangement, or partial conditioning. To this end, we present FlexCover, a cover generation model that conditions a pretrained text-to-song foundation model on a symbolic lead sheet. Our design enables alignment-free control: generated outputs preserve the tonal signature of the source without frame-level alignment, allowing flexible timing, structure, and segment-level generation. We further introduce a training curriculum with partially mismatched audio–symbolic pairs, improving diversity and robustness. We evaluate FlexCover with objective metrics, standard subjective ratings, and in-depth expert interviews that provide fine-grained diagnostic insights. Experiments show that FlexCover achieves state-of-the-art performance among open-source systems and is competitive with leading commercial models such as Suno-v5.5.