Home/ The Signal/ Industry/ AI Music Generation Legal Liability Isn't Being Decided by the Training-Data Fight
Licensing

AI Music Generation Legal Liability Isn't Being Decided by the Training-Data Fight

I asked a generator for a mid-tempo synth-pop bed — 104 BPM, F minor, warm analog pads, no vocals — and got back eight bars with a top line I could sing along to, because I had sung along to it…

A dimly lit professional recording studio at night, empty producer's chair pushed back from…

I asked a generator for a mid-tempo synth-pop bed — 104 BPM, F minor, warm analog pads, no vocals — and got back eight bars with a top line I could sing along to, because I had sung along to it before. Not a vibe match. Not "in the style of." A melody, sitting there in a 48kHz WAV with my prompt in the filename. I deleted it.

That render took forty seconds to make and would have taken a litigator's retainer to defend. It is also the part that most strategy decks about AI music generation legal liability leave out. Those decks are organized around training data: what went into the model, whether ingesting it was lawful, which appeal settles it. Meanwhile the thing that walked out of the machine and into a client's timeline was an output — and outputs are where enforcement is currently landing.

The myth: this all comes down to whether training was fair use

The version I hear most often, from publishers and from the tech side, goes like this. There are two futures. In one, courts hold that training a model on unlicensed catalog is fair use or falls under a text-and-data-mining exception, rightsholders lose their leverage, and licensing becomes a courtesy. In the other, training is held to require permission, and every platform has to come to the table. Everything turns on the input side. So the sensible move is to keep the licensing conversation warm, concede nothing, and wait for the appellate picture to clear.

It's a coherent myth. It is built on real cases, and the people repeating it are not naive. But it treats the training question as the only door into a defendant, and it isn't. It may not even be the one that opens first.

What the German ruling actually turned on

In late 2025, GEMA — the German collecting society, and one of the more litigious rights bodies in Europe — won a copyright case against Suno in a German court. The order included disclosure of relevant revenue and damages, which is the kind of remedy that gets read out loud in board meetings.

What matters more than the win is the reasoning. The court's analysis centered on what the service produced, not solely on what it consumed. The claim was framed around outputs that reproduced recognizable elements of protected works, and the court treated the apparent retention of substantial portions of those works inside the model as significant rather than incidental. That framing sidesteps the entire fair-use-in-training argument. You don't have to characterize ingestion at all if the finished render is the exhibit.

GEMA has run the same play before, bringing an action over song lyrics reproduced in a chatbot's responses. One society, two defendants, the same axis: show the output, match it to the catalog, argue the reproduction happened where the service was offered. It is worth reading that as a strategy rather than a coincidence.

Is making a song with an AI generator copyright infringement?

Usually no, sometimes yes, and the deciding factor is the output rather than the tool. Generating a track is not infringement because a model was trained on copyrighted music; it becomes an infringement problem when the specific render reproduces protected expression — a recognizable melody, a lyric, a distinctive hook — closely enough that an ordinary listener would connect the two. That test is the same one applied to human songwriters, which is why output claims travel so well: courts already know how to run them, and they don't require anyone to agree on what a model "is."

The corollary matters for anyone commissioning work: the liability attaches to the file you shipped, not to the vendor's training policy. A platform can have impeccable data provenance and still hand you a render that lands too close to something. A platform with murky provenance can hand you something clean.

Why output claims are structurally easier to bring

This is the mechanism, and it's worth being precise about, because it explains why the enforcement pattern is likely to continue regardless of how the training appeals resolve.

The evidence is already in your possession. An input-side claim requires you to establish what a private company put into a model — a fact that lives entirely inside the defendant's infrastructure and comes out, if at all, through contested discovery, protective orders, and expert fights over dataset reconstruction. An output-side claim requires a subscription, a prompt, and a hard drive. A plaintiff can build the evidentiary record before filing.

Memorization converts a technical artifact into an exhibit. Models that reproduce long passages of training material are, in the machine-learning literature, exhibiting a known failure mode. In a courtroom, that failure mode is a document. It also undercuts the most useful defense narrative — that the model learned abstract patterns rather than copies — without requiring the court to rule on training at all.

A wide overhead shot of a polished dark walnut conference table in a European…

Territory attaches at delivery. Training might happen on servers in a permissive jurisdiction. The render is delivered to a user sitting in Germany, or France, or wherever the service is offered. That gives a national court a hook on the act of communication to the public even when the model was built elsewhere. It is the reason "train somewhere friendly" is a weaker shield than it looks on a slide.

The remedies bite the product. Damages and disclosure orders aimed at output behavior land on the commercial service, not on a research process that already finished. You cannot un-train a model, but you can be ordered to stop shipping certain renders, and you can be ordered to say how much money the shipping made.

Input-side claim (training) Output-side claim (renders)
What you must prove Specific works were ingested and copied A specific output reproduces protected expression
Where the evidence lives Defendant's datasets and logs Your own hard drive
Jurisdictional hook Where training occurred Where the service is offered
Doctrinal state Contested, exception-dependent, evolving Long-settled substantial-similarity analysis
What a win produces Precedent about model-building Damages, disclosure, product-level orders

None of this makes output claims easy. Substantial similarity is a genuinely hard, expert-heavy, listener-dependent question, and it has produced inconsistent results in ordinary human songwriting disputes for decades. But it is hard in familiar ways, and familiar hardness is cheaper than novel hardness.

The rebuttal, and what stays unsettled

Suno's response has been that its technology transforms rather than reproduces, that the outputs are new works, and that US law — where its training took place — governs that question. The company has signaled it will appeal. That is a serious position, not a shrug, and the appeal has not been resolved as of writing.

It would also be dishonest to describe the underlying law as decided. In the US, the fair-use status of training on copyrighted material remains genuinely contested. Recent decisions involving book publishers and AI developers have turned on facts that do not map onto music: different markets, different acquisition practices, different output profiles. Nobody should be quoting a books-and-LLMs holding as though it governs a music generator, in either direction. And a German ruling on outputs is a German ruling on outputs — persuasive in the EU, informative in the UK and Japan, not binding anywhere else.

What has changed is the direction of travel. Enforcement is finding the path with the lowest evidentiary cost, and that path currently runs through the render rather than the corpus.

If you're deciding whether to license or to fight

Treat it as a portfolio question rather than a binary. A few things follow from the mechanism above.

Output-side enforcement is the cheapest leverage available to a rightsholder, which means it is also the fastest route to a licensing conversation that actually happens. Publishers who have concluded that litigation and licensing are alternatives are usually watching the expensive kind of litigation. The inexpensive kind — a documented set of outputs that land too close, brought in a jurisdiction where the service is sold — is a negotiating instrument.

If you're on the buy side, licensing AI-generated music into ads, games, or catalog, your exposure is contractual before it is doctrinal. Ask what the indemnity actually covers, whether it survives a third-party claim about output similarity rather than training, and whether the vendor documents provenance in a form you could hand to a client's legal team. We publish our own terms on the disclosure page for the same reason: a customer who has to guess is a customer with an unpriced risk. Any vendor that won't put it in writing has told you something.

And if you're building: the memorization problem is an engineering problem with a legal invoice attached. Output filtering, similarity screening before delivery, and honest logging of what a model produced for which prompt are cheaper than the disclosure order that asks the same questions later.

One thing to try this week

Pick twenty of your highest-value works — the ones that carry real synch revenue, not the long tail. Spend an afternoon running controlled prompts against two or three commercial generators: genre, era, tempo, instrumentation, and the kind of descriptive language a client would actually type. Save every render as an unedited WAV, log the exact prompt text, the date, the model version if the service exposes it, and a file hash. Then put an A&R ear on the results and mark anything that would make you nervous if a client shipped it.

That log won't win a case, and it isn't a substitute for counsel. What it will tell you, for the cost of one person's day, is whether you have an output problem at all — which is the fact you need before you can price a license, draft a demand, or decide that neither is worth the trouble.

The training-data fight will determine who owes whom five years from now; the output question determines who has leverage on Monday.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
N

Nova Reyes

Editor, The Signal

Nova Reyes edits The Signal and reviews AI music tools after a decade scoring indie games and short films; still owns four broken synthesizers. More by Nova Reyes →