REDACTED-Tunes: An Open 1.4M-Track Dataset and Perceptual Benchmark for AI-Generated Music
Robert Kaczmarczyk ⋅ TAWSIF AHMED ⋅ Felix Friedrich ⋅ Aidan C Erickson ⋅ Orian Sharoni ⋅ Dorien Herremans ⋅ Christoph Schuhmann
Abstract
Generative music platforms have reached commercial scale, with outputs approaching human-made quality. Yet music machine learning lags image-text research because no open music dataset matches the scale that catalyzed image-text foundation models: commercial recordings cannot be legally redistributed, and existing corpora are small, single-platform, or both. AI-generated music is the natural substrate to close this gap: We introduce $\textbf{REDACTED-Tunes}$, an open dataset of URLs and metadata for $\textbf{1{,}429{,}734 AI-generated music tracks}$ from Suno, Udio, and Mureka. The release ships public CDN URLs, 768-d audio and text embeddings, captions, ASR transcription embeddings (not raw transcripts), five-dimensional aesthetics scores, real and predicted engagement counts, and three-axis NSFW safety labels under Apache 2.0. We also release the models used to annotate the corpus: a 242M-parameter $\texttt{music-captioner}$ and fingerprint extractor, a $\texttt{SongEval}$-calibrated quality scorer, and an engagement predictor trained on platform play and upvote counts. To demonstrate downstream utility, we construct a perceptual benchmark from $\textbf{REDACTED-Tunes}$ and genre-matched human controls. In this study, $\textbf{61 participants annotated 591 song excerpts}$, and $\textbf{12 LLM-judge configurations}$ evaluated the same songs. Human listeners detect AI music above chance ($d' = 0.83$) but modestly (62.5\% accuracy), outperforming every LLM configuration on balanced accuracy. We further identify a quality--authenticity halo effect: songs receiving higher aesthetic ratings are more likely to be judged human-made; equivalently, songs judged real receive nearly two more aesthetic-quality points than songs judged AI-generated, independent of true provenance ($p < 10^{-26}$). All dataset artifacts, models, and code are available at https://anonymous.4open.science/r/anonymized-for-double-blind-review, alongside a live $\textbf{Tunes-Search demo}$ at https://anonymous.4open.science/w/anonymized-for-double-blind-review/search-tunes.
Chat is not available.
Successful Page Load