Skip to main content
Live
Main content

The Atlantic publishes searchable database of 21M+ songs used to train AI music models

Reporter Alex Reisner exposed four datasets, including two with 12M and 9M tracks, that Google and Stability have cited in research papers.

Jaeden Schafer
Editor in Chief · · 5 min read
The Atlantic publishes searchable database of 21M+ songs used to train AI music models

The Atlantic has published a searchable database exposing four music datasets used to train AI models, with the two largest containing 12 million and 9 million tracks. Reporter Alex Reisner identified the corpora and routed them through the magazine's AI Watchdog site, where anyone can now look up whether a specific song appears in training data. Two smaller sets in the collection still carry more than 100,000 songs each. The datasets have been downloaded thousands of times, and both Google and Stability AI have acknowledged using them in published research papers.

The disclosure widens an already active legal front. Music labels are suing AI music generators over training data, and the AI Watchdog index hands plaintiffs and individual artists a public tool to check which of their works sit inside the pipeline. Named tracks span Lady Gaga, Fred Again.., Radiohead, Aphex Twin, Wu-Tang Clan, Bruce Springsteen, and experimental composer Hainbach.

Millions of tracks are freely available in datasets, even if they're not supposed to be.
Terrence O'Brien, The Verge Weekend Editor

Not all of the data is pirated in the strict sense. The Free Music Archive dataset, for instance, is free to stream for personal use but requires licensing for commercial deployment — a line that a trained generative model selling subscriptions almost certainly crosses. Other corpora pull from sources where the terms are murkier or clearly violated by the collection method itself.

Key facts

  • 01The Atlantic's AI Watchdog now exposes four music training datasets totaling more than 21 million tracks, with the two largest holding 12M and 9M songs.
  • 02Two smaller datasets each contain over 100,000 songs, including material from the Free Music Archive that requires commercial licensing.
  • 03Google and Stability have both confirmed using the datasets in research papers, per reporter Alex Reisner.
  • 04Three of the four datasets are distributed as link lists pointing to YouTube and Spotify, requiring scraping tools that violate those platforms' terms of service.
  • 05Named artists in the data include Lady Gaga, Radiohead, Aphex Twin, Wu-Tang Clan, Bruce Springsteen, Fred Again.., and composer Hainbach.

The collection method is where the story gets sharper. Three of the four datasets Reisner found are not distributed as audio files at all. They are lists of links pointing at YouTube and Spotify, and turning the links into training data requires automated download tools that strip ads, bypass logins, and skip the monetization mechanisms that would otherwise pay creators.

Those tools violate the terms of service of both platforms. That detail matters because it shifts the legal question past fair-use arguments about training on copyrighted works and into straightforward platform-contract and anti-circumvention territory, which is harder to defend than a transformative-use claim.

Reisner laid out the mechanics directly in his reporting for The Atlantic.

Google confirmed using the data in research papers, which is a careful disclosure — research use generally enjoys broader legal latitude than commercial deployment, and Google has not said the corpora feed any shipping product. Stability AI's confirmation is more exposed given the company's commercial generative-audio ambitions and its existing litigation history over training-data provenance in image models.

The broader pattern is familiar. Large-scale training datasets in text, image, and now audio have repeatedly turned out to contain material their creators never licensed for AI use, surfaced months or years after the models built on them shipped. The web-scraping playbook that fueled the LLM boom is now being applied to music, and the same questions about consent, compensation, and platform terms of service are repeating beat for beat.

Related · from this week
Google's SynthID watermarking can weaken LLM safety guardrails, study finds
Jaeden Schafer · 5 min read →

What is new here is the searchability. Until now, most artists had no practical way to know whether their catalog had been scraped. The AI Watchdog tool changes the cost structure of a lawsuit: a musician or label can confirm inclusion in minutes rather than through discovery, which makes individual claims and class actions materially easier to file.

There are limits to what the database proves. Inclusion in a publicly available dataset does not by itself establish that a specific commercial model trained on a specific song, and AI developers have so far been reluctant to disclose training composition. Reisner notes it is impossible to know exactly who has used each set beyond the labs that have voluntarily said so.

Still, the disclosure tilts leverage toward rights holders heading into the next round of music-industry AI litigation. Generative-audio companies that have been raising at sharply higher valuations on the promise of licensed catalogs now have to explain, in court and in diligence, what is and is not in their training mix. The cheapest path forward — scrape YouTube, train, ship, apologize later — gets more expensive every time a tool like AI Watchdog goes live.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Anthropic logo
Security

Google's SynthID watermarking can weaken LLM safety guardrails, study finds

Lasso Security tested 6 open-weight models and found watermarking made some more likely to comply with harmful prompts under injection attacks.

Jaeden Schafer5 min read
Google logo
Security

Vendors move to block Google's data buy from bankrupt Spirit Airlines

Springshot says the auction hands Google 15 years of its IP; the EFF calls it the first public bankruptcy fight over employee data.

Jaeden Schafer5 min read
San Francisco orders Apple and Google to pull 13 AI 'nudify' apps
Security

San Francisco orders Apple and Google to pull 13 AI 'nudify' apps

City Attorney David Chiu says the two stores likely made millions from face-swap apps used to create nonconsensual nude images.

Jaeden Schafer5 min read