COMPANY
Suno
companytopic-notegenerative-audio
Overview
Suno is a US generative-audio AI company producing music-generation models trained on public-web audio corpora. The company entered the corpus on 2026-07-17 as the subject of a supply-chain compromise (reportedly traced to the Shai-Hulud npm worm) that exposed Suno source code and dataset manifests at a granularity discovery motions had not previously reached — establishing the first source-code-level provenance disclosure for generative audio. UMG partially settled with Suno in October 2025; Sony and residual UMG claims remain active in D. Mass. before Judge Saylor.
Timeline
- 2026-07-22-AI-Digest — A cyberattack on Suno leaked personal information of more than 55.3M users, per the Have I Been Pwned listing that surfaced today. Records include emails, phone numbers, addresses, and partial Stripe card data; per HIBP metadata the breach dates to November 2025, and leaked source code reveals prior scraping of Deezer, Genius, and YouTube. The disclosure gap is worth naming: Suno has not publicly confirmed the breach or notified users — the record count is HIBP’s, not Suno’s, and the discovery-to-disclosure lag is roughly eight months. Narrow read: the scale places it among the larger AI-SaaS breaches to date, and the leaked source revealing multi-platform scraping is a copyright-risk vector on top of the user-data one — extends the 2026-07-17-AI-Digest source-code-level provenance thread from what corpus data was ingested to what user data was retained. Structural read the corpus carries: Suno is a generative-AI startup that took in millions of user prompts, uploads, and account data as its core corpus — exactly the shape of vendor the “AI SaaS is now a high-value breach target” thread has been warning about. The HIBP-first disclosure pattern (breach detected by external researcher, not vendor) is the failure mode; audit AI-vendor incident-response commitments before piping user content in. 30-day watch: whether Suno issues a formal disclosure, and whether an amended discovery motion in the Sony case cites the leaked user records alongside the Jul 17 dataset manifests.
- 2026-07-17-AI-Digest — Suno source code and dataset manifests leaked in a supply-chain compromise reported to trace back to the Shai-Hulud npm worm, exposing a
youtube_musiccorpus of 2,013,545 clips / 113,879 hours plus tens of thousands of additional hours from Deezer, Genius, Pond5, IMSLP, Jamendo, and podcast RSS feeds. Prior music-AI provenance record (Udio’s April 2026 SDNY admission, The Atlantic’s earlier Suno/Udio corpus mapping) established the fact of YouTube scraping; today’s leak establishes it at the source-code level — dataset manifests with exact clip counts and hours per source, at a granularity discovery motions had not previously reached. Sony and residual UMG claims sit before Judge Saylor with dispositive motions currently reset to April 9, 2027 and statutory damages sought at up to $150K per work plus $2,500 per act of circumvention under DMCA §1201. Narrow read: the incremental legal risk is DMCA §1201 (circumvention of YouTube anti-scraping), not the pure infringement question the settled UMG matter mostly cleared. Structural read the corpus carries: training-set provenance for generative-audio labs is no longer an inference exercise — a source-code leak sets the discovery-motion template for the remaining Sony case and for the next round of publisher suits against any music-AI vendor with public-web-scraped training data. 30-day watch: whether Sony files an amended complaint that cites the leaked manifests as evidence — the fastest possible signal that source-code disclosures now materially move litigation.
Key Developments
- First Source-Code-Level Provenance Disclosure for Generative Audio (July 17, 2026): Supply-chain compromise (reported Shai-Hulud npm worm) leaks Suno source code and dataset manifests including
youtube_musicat 2,013,545 clips / 113,879 hours plus tens of thousands of hours from Deezer, Genius, Pond5, IMSLP, Jamendo, and podcast RSS feeds. Moves the music-AI provenance question from inferred (Udio’s SDNY admission, The Atlantic’s corpus mapping) to source-code-level, with exact clip counts and hours per source at a granularity that discovery motions had not previously reached. The incremental legal risk is DMCA §1201 (circumvention of YouTube anti-scraping) rather than the pure infringement question the settled UMG matter mostly cleared. Sets the discovery-motion template for the remaining Sony case and the next round of publisher suits against music-AI vendors with public-web-scraped training data.
Related
See also: MOC - Major Companies.