The Atlantic publishes searchable database of music in AI training datasets
Investigative reporting by The Atlantic uncovered four datasets containing millions of copyrighted tracks used to train AI models, now publicly searchable.
Last verified:
The Atlantic’s investigative reporter Alex Reisner has exposed a significant transparency gap in AI model training by uncovering and cataloging four datasets containing over 21 million musical recordings used to train commercial AI systems. According to The Verge, Reisner created a publicly searchable database documenting which artists’ work appears in these datasets, many without the artists’ knowledge or consent.
Scale and Composition of Discovered Datasets
The four datasets vary dramatically in size. Two datasets contain 12 million and 9 million tracks respectively, while the remaining two hold over 100,000 songs each. According to The Verge’s coverage of Reisner’s reporting, these datasets have been downloaded thousands of times across the industry. Major AI developers including Google and Stability AI have publicly confirmed their use of these training materials in published research papers, suggesting widespread adoption without proportional public disclosure.
The datasets span a broad cross-section of contemporary music. Artists whose work appears include mainstream pop acts like Lady Gaga and Fred again.., experimental musicians such as Aphex Twin and composer Hainbach, rock institutions including Radiohead and Bruce Springsteen, and hip-hop collectives like Wu-Tang Clan. The presence of such recognizable names underscores the datasets’ comprehensiveness and commercial value.
Technical Implementation and Legal Gray Area
Three of the four datasets operate as lists of URLs pointing to songs hosted on YouTube or Spotify rather than direct audio files. AI developers must download the actual audio using automated tools, which according to Reisner’s investigation, frequently bypass platform login requirements, advertisements, and revenue-sharing mechanisms. The Verge notes that these extraction tools directly violate the terms of service of both YouTube and Spotify, creating a tension between technical possibility and contractual obligation.
The Free Music Archive dataset, one of the sources, permits personal-use streaming without licensing but requires commercial licensing for derivative applications. Using such datasets for commercial AI model training likely exceeds the scope of personal-use licenses, creating potential copyright liability for model developers.
Why This Matters
The Atlantic’s searchable database addresses a critical asymmetry in AI training transparency. Most AI companies do not disclose which specific artists’ work trained their models, preventing creators from knowing whether their output influenced a system generating competing content. This database enables independent verification of training sources at a granular level—individual songs rather than aggregate dataset names. For recording artists, music labels, and policymakers evaluating AI regulation, the ability to search and identify whose work was used without permission raises urgent questions about fair compensation, attribution, and whether current copyright frameworks adequately protect creative workers in the generative AI era. The revelation that major vendors like Google and Stability openly used these datasets suggests industry normalization of a practice many musicians view as unauthorized use.
Frequently Asked Questions
Which music datasets were discovered?
The Atlantic identified four datasets: two containing 12 million and 9 million tracks respectively, and two smaller sets with over 100,000 songs each. Three of the four are distributed as lists of URLs linking to songs on YouTube or Spotify.
Did major AI companies use these datasets?
Yes, according to Alex Reisner's reporting, both Google and Stability AI have confirmed in published research papers that they used these datasets for training.
Are the datasets legal to use for AI training?
The datasets themselves are freely available online, but using them for commercial AI training may violate platform terms of service and copyright law. Many tracks require commercial licensing, particularly from sources like the Free Music Archive.
How can I search the database?
The Atlantic maintains an AI Watchdog site where the public can search through music, books, and other media included in these training datasets.