The Atlantic has published a searchable database exposing four music datasets used to train AI models, with the two largest containing 12 million and 9 million tracks. Reporter Alex Reisner identified the corpora and routed them through the magazine's AI Watchdog site, where anyone can now look up whether a specific song appears in training data. Two smaller sets in the collection still carry more than 100,000 songs each. The datasets have been downloaded thousands of times, and both Google and Stability AI have acknowledged using them in published research papers.
The disclosure widens an already active legal front. Music labels are suing AI music generators over training data, and the AI Watchdog index hands plaintiffs and individual artists a public tool to check which of their works sit inside the pipeline. Named tracks span Lady Gaga, Fred Again.., Radiohead, Aphex Twin, Wu-Tang Clan, Bruce Springsteen, and experimental composer Hainbach.
“Millions of tracks are freely available in datasets, even if they're not supposed to be.”— Terrence O'Brien, The Verge Weekend Editor
Not all of the data is pirated in the strict sense. The Free Music Archive dataset, for instance, is free to stream for personal use but requires licensing for commercial deployment — a line that a trained generative model selling subscriptions almost certainly crosses. Other corpora pull from sources where the terms are murkier or clearly violated by the collection method itself.
Key facts
- 01The Atlantic's AI Watchdog now exposes four music training datasets totaling more than 21 million tracks, with the two largest holding 12M and 9M songs.
- 02Two smaller datasets each contain over 100,000 songs, including material from the Free Music Archive that requires commercial licensing.
- 03Google and Stability have both confirmed using the datasets in research papers, per reporter Alex Reisner.
- 04Three of the four datasets are distributed as link lists pointing to YouTube and Spotify, requiring scraping tools that violate those platforms' terms of service.
- 05Named artists in the data include Lady Gaga, Radiohead, Aphex Twin, Wu-Tang Clan, Bruce Springsteen, Fred Again.., and composer Hainbach.
The collection method is where the story gets sharper. Three of the four datasets Reisner found are not distributed as audio files at all. They are lists of links pointing at YouTube and Spotify, and turning the links into training data requires automated download tools that strip ads, bypass logins, and skip the monetization mechanisms that would otherwise pay creators.
Those tools violate the terms of service of both platforms. That detail matters because it shifts the legal question past fair-use arguments about training on copyrighted works and into straightforward platform-contract and anti-circumvention territory, which is harder to defend than a transformative-use claim.
Reisner laid out the mechanics directly in his reporting for The Atlantic.
Google confirmed using the data in research papers, which is a careful disclosure — research use generally enjoys broader legal latitude than commercial deployment, and Google has not said the corpora feed any shipping product. Stability AI's confirmation is more exposed given the company's commercial generative-audio ambitions and its existing litigation history over training-data provenance in image models.
The broader pattern is familiar. Large-scale training datasets in text, image, and now audio have repeatedly turned out to contain material their creators never licensed for AI use, surfaced months or years after the models built on them shipped. The web-scraping playbook that fueled the LLM boom is now being applied to music, and the same questions about consent, compensation, and platform terms of service are repeating beat for beat.
What is new here is the searchability. Until now, most artists had no practical way to know whether their catalog had been scraped. The AI Watchdog tool changes the cost structure of a lawsuit: a musician or label can confirm inclusion in minutes rather than through discovery, which makes individual claims and class actions materially easier to file.
There are limits to what the database proves. Inclusion in a publicly available dataset does not by itself establish that a specific commercial model trained on a specific song, and AI developers have so far been reluctant to disclose training composition. Reisner notes it is impossible to know exactly who has used each set beyond the labs that have voluntarily said so.
Still, the disclosure tilts leverage toward rights holders heading into the next round of music-industry AI litigation. Generative-audio companies that have been raising at sharply higher valuations on the promise of licensed catalogs now have to explain, in court and in diligence, what is and is not in their training mix. The cheapest path forward — scrape YouTube, train, ship, apologize later — gets more expensive every time a tool like AI Watchdog goes live.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




