Home/ The Signal/ Industry/ Ninety Seconds: How AI Music Generation Actually Reaches Your Royalty Statement
Licensing

Ninety Seconds: How AI Music Generation Actually Reaches Your Royalty Statement

Ninety seconds. That was the render time. I fed a prompt into an AI music generation tool on a Tuesday afternoon — 92 BPM, brushed kit, a detuned Juno pad underneath, melancholy but not sad — and…

A dimly lit professional recording studio at dusk, empty except for a vintage analog…

Ninety seconds. That was the render time.

I fed a prompt into an AI music generation tool on a Tuesday afternoon — 92 BPM, brushed kit, a detuned Juno pad underneath, melancholy but not sad — and ninety seconds later I had two minutes of stereo audio occupying roughly the same emotional square footage as a cue I played keys on for an indie documentary in 2019. Not the same piece. Not close enough that any lawyer would take the call. Close enough that if the picture editor had needed a bed under a slow drone shot and had ninety seconds and a subscription instead of three days and a session fee, she would not have phoned me.

That number gets quoted in both directions. People selling these tools quote it as convenience. People who play for a living quote it as a threat. Both readings miss where the money actually moves, because the render is the fourth thing that happens, not the first. Three stages ran years before I typed anything. Two more run after the file exists, and those are the ones that touch your statement.

So here is the chain, in order, with the cost marked at each step.

First, your recordings become a dataset

Nothing here starts with a model. It starts with acquisition: someone assembles audio at a scale no A&R department has ever handled.

The methods are unglamorous. Bulk extraction from video platforms. Crawls of open-web audio. Stock and production-library catalogs whose terms of service were permissive enough, or vague enough, to be read as permission. Purchased or licensed catalogs, in the cases where a company decided to pay. In its 2024 response to the major labels' suit, one leading generator characterized its training material as, in effect, all the reasonably good-quality music accessible on the open internet — offered not as an admission but as a defense.

Notice the structural fact about this stage: it is the only one with no notification step. You get no email. There is no server log you are entitled to read, no line item, no ISRC-level report. If your playing is in there, you learned it the way the rest of us did — by inference, from an output that sounded familiar.

Session players and library composers are unusually exposed here, and it has nothing to do with quality. It is metadata. Work-for-hire cues frequently ship without performer credits. Library tracks are distributed widely, tagged exhaustively by mood and tempo, and sit on servers built to be crawled. A clean 48 kHz WAV labelled "tense ambient underscore, 70 BPM, no drums, no vocals" is close to an ideal training example: good audio, machine-readable description, and nobody whose name is attached loudly enough to make noise about it.

"Publicly available" and "licensed" are different words. Most of the argument of the next decade lives in the gap between them.

Second, the model learns the shape of your playing, not the file

What happens next is compression, then prediction. An encoder turns waveforms into discrete tokens — a learned codebook of audio fragments. A transformer learns to predict which token follows which, conditioned on text. A decoder turns predicted tokens back into audio.

Which means the honest version of the industry's favorite defense is true: your file is not sitting in the weights in any recoverable form. Nothing is retrieved. "We don't store your song" is technically accurate and almost entirely beside the point, because what does survive is the relationship between features. The drag you put on the backbeat. The way you voice a minor ninth with the third on top. The particular dead thud of the room you tracked in. None of that persists as your performance. It persists as a slightly raised probability that those characteristics occur together when someone types a mood.

You will hear the counter-argument constantly: human musicians learn the same way. That is true as far as it goes, and it goes about half the distance. A human learner is inside the economy they learned from — buys the record, takes the lesson, gets hired, gets heard, gets older and teaches somebody else. A weights file does none of that. The difference worth arguing is scale and reciprocity, not mechanism, and musicians who argue the mechanism tend to lose, because on the mechanism the other side is partly right.

One caveat cuts the other way. Models do sometimes memorize. Researchers and litigants have coaxed outputs that reproduce recognizable melodies and lyric sequences close to verbatim, usually with pointed prompting. It is not the normal behavior. It exists, it has been demonstrated in filings, and it complicates the tidy claim that nothing is ever copied.

Is AI-generated music copyright infringement?

It is at least three separate legal questions, and as of writing they are resolving in different directions. First, the training copy: whether reproducing recordings in order to train a model is infringement, fair use, or covered by a text-and-data-mining exception. That is genuinely unresolved and depends heavily on jurisdiction. Second, the output: a generated track infringes only if it is substantially similar to a specific protected work. Genre, tempo, production aesthetic, and "sounds like that band" are not protected. Third, the voice: identifiable vocal likeness is being handled mostly outside copyright, through right-of-publicity and new likeness statutes rather than through infringement claims.

A close profile portrait of a session musician standing alone in a soundproofed live…

The practical consequence is uncomfortable. A track that sounds exactly like your sound while copying none of your notes is usually not a copyright problem at all. That is not a loophole anyone engineered. It is how the law has always worked, and it was survivable for a century because imitating you convincingly used to require hiring someone approximately as good as you.

What you hold What it actually covers What it does not stop
Composition (ISWC) melody, lyrics, specific structure a new melody built in your idiom
Sound recording (ISRC) that master, that take a re-performance of the same feel
Voice / likeness identifiable voice, where statute exists a timbre trained to sit next to yours
genre, tempo, arrangement habits, mix character

This describes terrain, not outcomes. Terms differ by territory and by contract, and the case law is moving; if you are deciding something with money attached, read your agreement and ask a lawyer who practices where you file.

Third, the prompt collapses everything toward the median

Now the ninety seconds.

The output is worse than the marketing and better than the dismissals. Cymbals arrive as a smear of compressed noise rather than a strike with a decay. Transients get rounded, so a snare lands as a shape instead of an event. Sustained vowels wander in formant, which is why vocals remain the hardest thing these systems do. Arrangements reach the chorus on schedule without earning it. Loop points breathe, which matters if you are building adaptive game audio and do not want an audible seam every 32 bars. And prompt roulette is real: eight renders for one usable result is a normal afternoon, and the usable one is usable largely because it is unremarkable.

That last part is the problem, not the consolation. These systems are trained to produce the likely continuation, and the likely continuation is the center of the distribution. The center is not where the records people love live. It is where the beds people need live — underscore, corporate video, podcast intro, menu loop, the two minutes under the b-roll that nobody will ever name or search for.

That market is not the embarrassing part of a career. For a great many working musicians it is the part that pays rent between the projects that pay pride. The tools are weakest exactly where music is most distinctive and strongest exactly where music is most functional, and functional is the base of the pyramid the rest of it stands on.

Fourth, the upload, where aesthetics become arithmetic

A render sitting on a hard drive costs you nothing. The financial mechanism starts when it is distributed.

Most streaming services pay from a pooled model: subscription and ad revenue for a territory goes into a pot, and rights holders are paid according to their share of total qualifying streams. Your per-stream rate is not a price anyone set. It is a quotient. Anything that inflates the denominator reduces your number, and it does not need to be good, or heard by many people, to do so — it needs to exist and to accumulate qualifying plays.

That is the whole mechanism. It explains why catalog flooding works as a business even when the individual tracks are mediocre, and why services that built their own detection have reported that fully AI-generated material now accounts for a double-digit share of daily uploads, a figure at least one of them has revised upward more than once.

Three amplifiers sit on top of it. Functional playlists — focus, sleep, ambient, lo-fi — are substitution-friendly by design, because the listener's requirement is a mood at a volume rather than a specific artist. Minimum-stream thresholds, introduced by several services to demonetize the long tail, land on the composer with forty tracks rather than on the operation uploading four hundred a week. And soundalike releases, published under names adjacent to real artists and occasionally attached to dead ones, siphon intent-driven search traffic that was aimed at a human being.

None of this requires anyone to have infringed anything. It requires only that the pool be finite and the supply be cheap.

Fifth, why almost nothing catches it

The detection layer was built for a different crime.

Acoustic fingerprinting — the technology behind Content ID and the app that names the song in a bar — matches a candidate against reference recordings. It is very good at finding your master inside somebody else's upload. It has nothing to say about a track that reproduces your voicings, your tempo, your reverb tail and your habits without containing a single sample of your audio, because there is no reference to match. The system is answering a question about copies while the actual event is a question about style.

An overhead flat-lay of a cluttered archival desk in a windowless room, hundreds of…

The newer tools are partial. Classifiers that flag likely AI-generated audio do work above chance, and they are degraded by exactly the things a determined uploader does anyway: transcoding, a light re-master, running the file through analog gear, or having a human overdub eight bars. Watermarking helps only if the generator embedded one, and survives only if nobody transcodes it away. Content-provenance metadata such as C2PA is genuinely useful at the point of creation and tends to be stripped at the first re-upload, because most platforms rewrite files on ingest. Disclosure fields on distribution forms are, as of writing, largely voluntary and self-reported.

Stack those and you get the outcome that matters: the burden of noticing, proving, and objecting falls on the individual performer, who has the least time, the smallest legal budget, and no view into stage one.

Sixth, what is actually moving

Three things, at different speeds.

Litigation is the loudest. The major labels sued the two best-known generators in 2024; publishers have brought their own actions over lyrics; similar cases are running in several jurisdictions. By the time you read this the posture will have changed, possibly more than once. Watch what the resolutions look like rather than who filed, because there is a specific risk in how these tend to end: a suit between a large rights holder and a platform can settle into a licensing arrangement in which both parties do well and the performer on the recording is not a party to the deal at all. If your work is inside a label or library catalog, the negotiation that decides whether you are training data may already be happening without you in the room.

Legislation is slower and, for once, moving in a useful direction. The EU's AI Act obliges providers of general-purpose models to publish a sufficiently detailed summary of training content. Several jurisdictions have passed or proposed voice-and-likeness protections — Tennessee's 2024 statute on voice being the most cited — that treat vocal identity as a right independent of copyright. Transparency and opt-out regimes are under discussion in multiple parliaments, with the usual gap between passing a rule and enforcing one.

Collective licensing is the quietest and possibly the most consequential. Collecting societies and rights databases are the only infrastructure that has ever successfully metered a use case this diffuse. They did it for radio. They could do it for training corpora, but only if the corpora are disclosed.

Which is why disclosure is the fight worth your energy, ahead of the arguments about whether machines can be creative. You cannot license what you cannot see, and you cannot price what nobody has to count.

What to do this quarter

None of the above is actionable on a Tuesday. These are.

  • Fix your metadata before you argue about rights. ISRC on every master, ISWC on every composition, splits registered, performer credits delivered. An unattributed recording is both easier to ingest and harder to claim.
  • Read the AI clause in every contract you sign this year. Look specifically for language granting the right to use the recording for machine-learning training or to create derivative models. Ask for the carve-out; get it in writing; expect pushback from buy-out clients.
  • Check your library's terms of service, not just your original agreement. A number of catalogs quietly updated their licensing language to permit data-licensing deals. Your old contract may point at their current terms.
  • Keep provenance you control. Dated session files, take sheets, raw stems, project archives. Embedded provenance metadata strips on upload; your own archive does not.
  • Set alerts on your artist names and your own name across DSPs monthly. Releases you did not make are easier to remove in week one than in year one.
  • Price and sell what the pipeline cannot do. Revision under notes, sync to a moving picture lock, live performance, being reachable at four o'clock on Thursday when the edit changes.
  • Know what the tools actually license. Ownership, commercial use, and indemnity vary sharply between generators and change without much announcement; we keep a running read on how each one's terms are written, and it is worth checking before you or a client builds a deliverable on one.

Ninety seconds, reconsidered

I still have the file. It is fine. The pad is a little wide, the brushes sound like brushes described rather than played, and at 1:14 there is a swell that arrives from nowhere and resolves into nothing. If a client had asked, I would have taken it out.

But the number was never a statement about how fast a machine makes music. Stage one took years and produced no receipt. Stage two turned a few hundred thousand hours of human decision-making into a probability surface. Stage four dropped the result into a pool that pays you by division. Stage five made sure nobody had to tell you. The render is quick because every expensive part of it was completed in advance, by people who were not asked and have not been counted.

Ninety seconds is not the speed of a machine making music. It is the speed of a machine spending something it never had to buy.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
T

Theo Brandt

Tutorials Writer

Theo Brandt writes step-by-step tutorials for AI music tools — prompting, stem workflows, and release prep — from a bedroom studio that started with a cracked DAW and a $60 mic. More by Theo Brandt →