The Atlantic created a searchable database of the music used to train AI

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- The Atlantic released a searchable database built by reporter Alex Reisner, cataloging four music datasets used to train AI models and exposing their contents to the public.
- Two of the datasets are massive at 12 million and 9 million tracks, while the other two each contain over 100,000 songs — and all have been downloaded thousands of times.
- Google and Stability have confirmed in research papers that they used these datasets to train their AI models.
- Three of the datasets are distributed as lists of links to YouTube or Spotify songs, with automated download tools that bypass platform logins, ads, and revenue mechanisms in violation of terms of service.
- Artists named in the datasets range from Lady Gaga, Radiohead, Aphex Twin, Wu-Tang Clan, and Bruce Springsteen to experimental composer Hainbach.
- Some sources, like the Free Music Archive, are free for personal streaming but require licensing for commercial use — a boundary that AI training has apparently crossed without permission.
Why it matters: Google and Stability are confirmed users of datasets that aggregate tens of millions of tracks via tools violating YouTube and Spotify's terms of service, putting both companies in legal and ethical crosshairs with major-label artists. The Atlantic's decision to make the data fully searchable means any artist, label, or regulator can now instantly check whether their catalog has been ingested for AI training — a transparency escalation that could accelerate licensing demands or lawsuits.
Ask SkimNews



