Home/ The Signal/ Industry/ The Paper Trail: Three Ways to Find Out What an AI Music Generation Model Was Trained On
Training Data

The Paper Trail: Three Ways to Find Out What an AI Music Generation Model Was Trained On

The first hard evidence in the Suno case was not a document. It was audio. When the major labels sued in June 2024 — Universal, Sony and Warner, coordinated through the RIAA, Suno drawing the District…

A photorealistic overhead photograph of a heavy stack of legal documents spread across a…

The first hard evidence in the Suno case was not a document. It was audio.

When the major labels sued in June 2024 — Universal, Sony and Warner, coordinated through the RIAA, Suno drawing the District of Massachusetts and Udio the Southern District of New York — the complaints arrived with outputs attached. Generations the plaintiffs said tracked specific masters closely enough that a judge could hear the resemblance without an expert standing by to narrate the waveform. Not a leaked manifest. Not a config file. Renders. Somebody typed a prompt, listened, and thought: that is ours.

That is where the training-data question actually sits for anyone responsible for a catalog. The companies building AI music generation tools have answered "what did you train on" in three separate registers — what they volunteer, what a docket pulls out of them, and what the model concedes when you interrogate it directly. Three trails, three species of fact, arriving at wildly different speeds with wildly different odds of surviving contact with a courtroom or a term sheet.

They get treated as interchangeable. They are not. Most of the noise in this area comes from grading a press statement, a sworn pleading and a breach rumour on the same scale.

What we are grading, and on what

Five criteria, applied to each trail:

Specificity. Does it name actual sources — catalogs, platforms, domains, individual recordings — or does it stay at the altitude of "publicly available data"?

Verifiability. Can a third party reproduce the finding without taking the reporter's word for it?

Speed. How long after the training run does the information surface?

Legal usability. Does it hold as evidence, or underpin a representation and warranty in a licence?

Durability. Is it still true after the next retrain?

Criterion Voluntary disclosure The litigation record Adversarial testing
Specificity Category-level only High, where it survives sealing Work-by-work, one at a time
Verifiability None — self-reported High — sworn, sanctionable High — anyone can rerun it
Speed On the company's clock Slow; a year-plus to real production Same afternoon
Legal usability Mainly as an admission The gold standard Needs an expert to carry it
Durability Decays every model version Frozen at the discovery cut-off Snapshot; re-run per version

Trail one: what they say on purpose

Voluntary disclosure in this sector has converged on one sentence with three interchangeable phrasings: publicly available data, licensed and public sources, material accessible on the open internet. That is a category, not a list. It tells a rights manager nothing about whether a particular catalog was ingested, which is the only question that moves an exposure model.

The interesting exception is when a company says something specific enough to be used against it. Suno's answer in the Massachusetts case, filed in August 2024, described training on essentially all music files of reasonable quality accessible on the open internet, and framed the whole exercise as fair use. That is not disclosure in the transparency sense. It is a legal position. It is also the most useful voluntary statement anyone in this sector has produced, precisely because it went in under signature and cannot be quietly revised in a footer next quarter.

Everything else in this tier decays. Model cards get rewritten. Blog posts get edited without changelogs. Terms of service get amended on notice periods measured in days. A statement about version 3 tells you nothing about version 4.5, and these systems retrain on a cadence that outruns your diligence file.

The one thing voluntary disclosure is genuinely good for is dating. What a company was willing to claim in public at a given moment is worth having in a folder, timestamped, because the claim will change and the earlier version becomes leverage. Archive the page. Do not rely on the vendor's copy.

Grade: specificity low, verifiability zero, speed on their clock, legal usability confined to admissions, durability poor.

Trail two: what the docket drags out

This is the only trail that produces sworn specificity. Interrogatories, requests for production, a deposition of whoever actually built the ingestion pipeline, expert reports reconstructing a dataset from server logs and repository history. In an ordinary infringement case, discovery is exactly where "publicly available data" becomes a list of domains, a set of scraper configurations, and a date range you can check against your own release schedule.

The problem is structural. This trail is built to close, and it closes three ways.

Protective orders. The material worth having comes in attorneys'-eyes-only. A favourable ruling can land with the dataset description redacted to the point of uselessness for anyone not already inside the case.

A close-up photorealistic photograph of a professional recording studio mixing console in a darkened…

Settlement. Through late 2025 the posture shifted from pure litigation toward settlement-plus-licensing, with announced arrangements between major rightsholders and both named defendants; terms were not fully public. As of writing, that is the direction of travel. A settlement that converts a copyright claim into a commercial partnership also converts the training-data question from a public evidentiary issue into a private contractual one. The list may well exist in detail. It exists inside a deal you either signed or did not.

Mootness. If a court reaches the four factors and lands on fair use for the ingestion step, granular provenance may never be adjudicated at all. The question that keeps your department awake gets resolved without ever being answered.

There are parallel fronts, and they matter for exactly this reason. GEMA has pursued Suno through the German courts, and the European proceedings have shown appetite for treating output resemblance as a question distinct from training legality. Different jurisdiction, different disclosure culture, different sealing practice. Following more than one docket is not thoroughness. It is the only way to get more than one draw at the same record.

One underrated asset in this tier: the pleadings themselves. A complaint filed by a well-resourced plaintiff contains that plaintiff's best available reconstruction of what happened, assembled by people with subpoena power and a forensic budget. It is advocacy, and it is also the cheapest research you will ever read.

Grade: specificity high where it survives, verifiability high, speed slow, legal usability unmatched, durability frozen at the discovery cut-off.

Trail three: what the model concedes

This trail has moved fastest and gets the least respect from people who work primarily in documents.

Memorization probes. Generative systems trained on enough copies of a work will, under the right conditions, reproduce recognisable fragments of it. Not the whole master — a voicing, a phrase, a rhythmic signature that has no business surviving a stochastic process. This is the mechanism behind the output exhibits in the complaints, and it is available to anyone with a paid account and a quiet afternoon.

Membership inference. The broader research family asks whether a specific item was in the training set by measuring how the model behaves around it — confidence, loss, the shape of the output distribution. Applied to audio it is noisier than it is on text, and results need an expert to interpret responsibly. It is also the only technique that scales toward a portfolio-level answer rather than a track-level one.

Your own infrastructure. If you host catalog anywhere — a promo portal, a sync library, a label site with streamable previews — your access logs, CDN records and robots.txt history are evidence you already own, sitting in a bucket, uncollected. Crawler traffic against a directive is a documented fact with a timestamp on it. Rights managers routinely commission expensive outside research while the relevant logs age out of retention at ninety days.

Leaked artifacts. The breach category. Reporting on code said to have surfaced from a Suno security incident described scraping infrastructure pointed at large public music, lyrics and video platforms — broadly the names anyone in the room would have guessed. The company's line was that the code was stale and out of service, and that user data was not at risk. Both statements can be true. Neither is verifiable from outside. As evidence this material is close to worthless: no chain of custody, no authentication, no reliable date. As a research lead it is excellent, because it tells you which platforms to probe first.

That distinction is the correct filing system for every leak story in this sector. Worthless as proof. Valuable as a map.

Here is the shape of a probe that is worth running. Deliberately generic, deliberately unremarkable:

An uptempo 1970s Philadelphia soul arrangement: sweeping strings,
four-on-the-floor kick, male tenor lead, tambourine on the backbeat.
118 BPM, key of E-flat.

You are not asking the model to copy anything, which matters both methodologically and for how the exercise looks later. You are testing whether a bland genre description collapses toward one specific arrangement. Reseed it twenty times. Genuine generalisation drifts — different string voicings, different drum feel, tempo wandering a few BPM either side. Memorisation converges.

Do AI music companies have to disclose their training data?

In the United States, as of writing, there is no general federal statute requiring an AI developer to publish what it trained on. The obligations that exist are regional and recent. The EU AI Act requires providers of general-purpose models to publish a sufficiently detailed summary of training content, with obligations for general-purpose models phasing in from August 2025 and a template issued through the AI Office. California's AB 2013 requires developers of generative systems made available to Californians to post documentation about the datasets used, on a compliance schedule running into 2026. Neither instrument produces a track listing, and enforcement of both is largely untested.

A photorealistic portrait of a person seated alone in a dim audio listening room…

What a "sufficiently detailed summary" realistically gets you is narrative-level: the major public datasets by name, the categories of scraped material, the broad provenance buckets. That is enough to establish whether a provider scraped a platform that hosted your catalog. It is not enough to establish that a specific master was in the set, and it is certainly not enough to compute a damages figure.

Which is the honest read on the regulatory tier: it is the only trail engineered to be public and durable, and it is the weakest of the three on specificity. It gives you a place to start an argument, not a place to finish one. Treat the filings as free discovery you did not have to litigate for, and read them against what the same company said in its marketing.

What memorisation actually sounds like

I cannot read a docket the way your associate can. What I can do, after a decade of scoring games and shorts and living inside other people's mixes, is tell you when a render has stopped generating and started remembering. Five tells, in rough order of how hard they are to explain away:

  • Tempo lock. A generic prompt with no BPM given, and the outputs keep landing on the source recording's exact tempo. Models have tempo priors; they do not have this one for no reason.
  • Arrangement collapse. The same horn voicing, the same fill in the same bar, the same drop on the same count, across independent reseeds. Variation is what these systems are for. Its absence is a finding.
  • Room as fingerprint. Reverb tails and pre-delay are production decisions made in a specific room on a specific day. When a render reproduces a plate tail you recognise from a 1978 master, the model did not derive that from a genre description.
  • Ad-libs and tags. A vocal flourish, a producer tag, a spoken syllable that belongs to one artist. Hardest of all to attribute to coincidence, and the reason vocals stay the sharpest edge of this litigation.
  • Key gravity. Outputs clustering in one key when the prompt named none, and it happens to be the key of the obvious reference.

Document every run: model version, exact prompt string, seed where the interface exposes one, wall-clock timestamp, and the audio file itself with its hash. Then push the outputs through an audio fingerprinting service and keep the report. Undocumented probing is a hobby. Documented probing is an exhibit in waiting — and it is also the only thing in this whole landscape that runs on your schedule instead of somebody else's.

Seven things to ask for, in writing

We put the same questions to every tool we cover at City of Punk, and the pattern of what gets answered is itself informative. For a diligence file or a licence negotiation:

  1. The model version and training cut-off date the representation actually covers — not "our models," a version string and a date.
  2. Named sources, not categories. "Licensed and publicly available" is a refusal dressed as an answer.
  3. Removal on retraining, not filtering at output. Output filters are cheap; retraining is expensive. A promise to block is not a promise to forget, and the two get deliberately blurred in redlines.
  4. An audit right exercisable by a named third party under NDA, with a defined scope and a response window.
  5. Indemnity scope. Outputs only, or ingestion too? What is the cap, and does it survive a class action?
  6. Notification on retrain, so your representation does not silently expire the next time they ship.
  7. Their EU summary and any California posting attached as a contract exhibit — which converts a regulatory statement into a warranty you can sue on.

Item three is where most negotiations quietly go sideways, and item seven is the cheapest leverage on the list.

Where this leaves you

Voluntary disclosure is worth keeping and worth nothing on its own — a dated admission, useful mostly when the story changes. The litigation record is authoritative and self-extinguishing: it produces the specificity everyone wants, then seals it, settles it, or moots it, and each quarter of settlement activity makes a public answer less likely rather than more. Adversarial testing proves presence and can never prove absence, needs an expert to walk it into a courtroom, and expires with every retrain.

And it is the only one of the three that runs when you decide it runs.

The industry is standing around waiting for a disclosure event — a court order, a regulatory filing, a leak that finally names the sources. That event may not arrive in a form you can use, and the deals being signed make it less probable every quarter. The evidence you can generate this week costs a subscription, a reference library you already own, and someone with ears.

Nobody is going to hand you the training list — but you own the reference recordings, and the model will tell you which ones it already knows.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
J

Juno Park

Game Audio Writer

Juno Park covers AI sound design and game audio workflows — foley, loops, and middleware — after seven years cutting assets for mobile and indie titles. More by Juno Park →