Doctorcrypto About RSS Subscribe
Doctorcrypto
HomeBusiness › Leaks Reveal Suno Fed Thousands of Hours of Deezer, YouTube and Pond5 Data Into Its AI
Business

Leaks Reveal Suno Fed Thousands of Hours of Deezer, YouTube and Pond5 Data Into Its AI

By Diego Whitfield · · 2 min read

Newly leaked source code appears to expose how AI music generator Suno assembled its training library, pointing to thousands of hours of audio sourced from platforms including Deezer, YouTube, and stock media provider Pond5.

What the Leaks Reveal

According to the leaked material, Suno's training pipeline drew heavily on content pulled from major streaming and media services. The documents suggest that substantial portions of the company's dataset were built from music and audio harvested across these platforms, raising fresh questions about how the popular tool was trained.

The revelations offer a rare glimpse into the inner workings of an AI music company that has largely kept its data practices under wraps. Suno, like many generative AI firms, has faced persistent scrutiny over the origins of the material used to teach its models to produce original-sounding tracks.

The leak pulls back the curtain on a training process the company has never fully disclosed.

Why It Matters

The disclosures land amid an ongoing debate over the legality and ethics of training generative AI on copyrighted works. Music rights holders and platforms have increasingly pushed back against companies that scrape or ingest their catalogs without explicit permission or licensing agreements.

For Suno, the leaked details could intensify legal and reputational pressure. The company has been among the most prominent players in AI-generated music, and evidence that its dataset relied on content from established services may fuel arguments from artists and labels who claim their work was used without consent.

The sources reportedly implicated span a broad range of content types:

  • Streaming catalog audio from Deezer
  • Video and music content from YouTube
  • Stock media from Pond5

The Bigger Picture

The situation reflects a wider reckoning across the generative AI industry, where questions about training data have become a central battleground. As lawsuits and licensing disputes mount, the pressure on companies to prove the provenance of their datasets is only growing.

Whether these leaks translate into formal legal action remains to be seen, but they add to a mounting body of scrutiny facing AI music tools and the data practices that power them.

Was this useful?👍 Yes👎 No