Home/ The Signal/ Industry/ Scrape, Caption, Train, Ship: How AI Music Generation Actually Gets Made
Suno

Scrape, Caption, Train, Ship: How AI Music Generation Actually Gets Made

The render came back at 128 BPM, minor key, with exactly the sludgy 808 I had asked for. About eleven seconds in, buried in the tail of a reversed cymbal, there was a syllable.

A dim home studio at night photographed at eye level from just behind an…

The render came back at 128 BPM, minor key, with exactly the sludgy 808 I had asked for. About eleven seconds in, buried in the tail of a reversed cymbal, there was a syllable. Not a word — a shape, the way a producer tag sounds through a wall in the next apartment. I bounced the same prompt four more times and never heard it again.

That is the uncomfortable part of AI music generation: you cannot audit a render. You can listen to it, and you can wonder what is underneath it. For most of the last two years, the answer to what is underneath it came from lawyers — filings, motions, carefully-worded blog posts. Then a security breach at Suno put a different kind of document on the table, and the interesting thing about the leak was not that it happened. It was that it matched what the company had already said in court.

So it is worth walking the machine in order. Not the marketing order — prompt in, song out — but the actual causal order. What happens to a piece of music first, what happens to it next, and what happens to it last, when it lands on your drive with a license attached and a video edit due Friday.

First: the collection

The pipeline starts with acquisition, and acquisition is the step with the fewest moving parts and the most legal exposure.

When a music-model company says it trained on publicly available audio, the practical meaning is narrower and stranger than it sounds. Scraping YouTube means pulling the audio track out of video containers at scale. What comes back is not a label catalog. It is live sets, DJ mixes, karaoke uploads, production-library cues, stems that a producer posted for a remix contest, full commercial records uploaded by people who did not make them, and a lot of audio nobody would call a song. Suno's own filings had already conceded a training pool that amounted, more or less, to whatever listenable audio the open web would surrender.

After the breach, reporters at 404 Media went through the exposed material and published counts: roughly two million clips, and by one tally well north of a hundred thousand hours of audio. Treat those as the figures available at the time rather than a final inventory. Even taken loosely, they describe more than a decade of continuous sound, collected without any mechanism for the people who made it to know they were in it.

Next: the grinder

Here is the step almost nobody covers, and it is the one that decides everything downstream.

Once audio enters a training pipeline, it stops being a song. It gets transcoded to a uniform sample rate, sliced into clips — a few seconds to half a minute is typical — and loudness-normalized so nothing dominates for the wrong reason. Then each clip is run through an audio-captioning model that writes a text description of what it hears: male vocal, lo-fi hip hop, vinyl crackle, roughly 92 BPM, warm low end. That caption-and-clip pair is the training example. Everything else is overhead.

Notice what falls off the truck. The artist's name was in a filename or an ID3 tag, and the filename is not what the model learns from. By the time training begins, the thing that would tell you whose performance is in the batch has already been discarded as metadata.

This is why "take my music out of the model" is not a button anyone can build cheaply. Removal implies a lookup table from work to weights, and the architecture never keeps one. Opt-outs offered after the fact usually mean we will exclude you from the next training run, which is a different promise wearing the same coat.

What the hack actually revealed

A security breach does not normally explain how a model was trained. This one did, and the reason is worth stating plainly: the exposed material corroborated claims the company had already made in its own defense. The value of the leak was not novelty. It was confirmation, in file listings and durations, of a scale that had previously existed only as a legal characterization.

That distinction matters if you are reading this as a tech story rather than a music one. The breach did not catch anyone lying. It removed the abstraction. "Publicly available material" is a phrase you can argue with. A directory of ripped audio measured in six figures of hours is a thing you can count, and counting changes how a jury, a legislator, or a working musician reacts to the same underlying fact.

The reaction from musicians in the weeks after was less about surprise than about confirmation fatigue — the sense of having been told for two years that the question was complicated, then seeing the file list.

Then: training, and why the output is not a copy

The model that comes out the far end does not contain WAV files. It contains weights: a very large set of learned relationships between caption tokens and compressed audio representations. Generation samples from those relationships. Nothing is retrieved.

A cavernous server-room aisle shot with a wide 24mm lens from floor height, endless…

This is technically true, and it is the load-bearing beam of the industry's defense. Verbatim reproduction does happen, but it happens as memorization — the artifact of material that appeared often enough, in similar enough form, that the model overfit to it. That is also why my ghost syllable never came back. Sampling temperature moved, and the neighborhood the model wandered into moved with it.

The plaintiffs' answer is not that the model is a jukebox. It is that the copying already happened, at ingestion, before any weights existed. Both statements can be accurate at once. That is precisely why the litigation is hard, and why arguments about whether a generated track "sounds like" a specific record tend to miss the claim being litigated.

Worth saying, since this publication is not in the business of selling the technology: the outputs still have obvious ceilings. Vocals smear on sustained notes. Dense arrangements go mushy in the low mids. Prompt roulette is real, and getting a usable 30-second bed can take a dozen bounces.

Last: the legal argument, which arrives after the model

Causally, this step comes last, and that ordering is the whole story. The dataset was assembled, the model was trained, the product shipped, and only then did the question of permission get formally adjudicated.

Two US federal rulings in 2025 involving models trained on books are cited constantly now, usually as though they settled something. They did not settle this. One found the training itself sufficiently transformative to sit inside fair use while treating the sourcing of the underlying library as a separate problem the defendant did not get to skip. The other went the defendant's way largely because the plaintiffs failed to build a record on market harm — a procedural outcome, not a blessing.

Music may not map onto those cases cleanly, for two structural reasons.

First, a recorded track carries two copyrights: the composition and the sound recording. Training on the audio touches both, and the case law around sampling sound recordings has historically been less forgiving than the case law around quoting text.

Second, and sharper: market substitution. A model trained on books produces prose that competes with books somewhere in the general economy of reading. A model trained on production-library cues produces production-library cues, sold to the same video editor, for the same 90-second explainer, at a lower price. The fourth fair-use factor asks about harm to the market for the original work. It is difficult to think of a cleaner example of the input's market being the output's market.

None of that predicts an outcome. It describes why the text-training precedents may be a weaker shield here than the confident citations suggest.

What lands on your drive

Which brings the mechanism to you — the game developer who needs adaptive loops, the editor who needs a bed under a voiceover, the podcaster tired of the same three intros.

Before you ship a generated track in anything commercial, check these six things in the vendor's current terms:

  • Which tier actually grants commercial rights. On most platforms it is paid-only, and the free tier is explicitly non-commercial.
  • Whether those rights survive cancellation. Some grants are tied to an active subscription. If yours is, a lapsed card can affect work already shipped.
  • Exclusivity. You almost certainly do not have it. The same prompt can hand someone else a very close neighbor of your cue.
  • Deliverables. 48 kHz WAV or MP3 only, and stems or a mixed two-track. Without stems you cannot duck an element under dialogue; you can only duck the whole bed, and it will sound like it.
  • Indemnification. Does the vendor defend you if a claim lands on your video, or does the license grant rights while leaving the risk entirely with you.
  • Platform registration. Content ID matches audio fingerprints, not ownership. Understand whether you can register a generated track, and what happens if someone else registers one that resembles yours.

Then do the boring thing: keep the prompt, the seed if the tool exposes one, the generation date, and a copy of the license text in the project folder. If a claim ever arrives, the ability to show what you generated and when is worth more than any assurance in a marketing page. We track how these terms differ across the major generators here at City of Punk because they change without announcement — read the vendor's own page before you commit anything that matters. And none of the above is legal advice; it is a list of places where the terms tend to bite.

The myth is that AI music generation is a machine that learned music the way a student does — listening broadly, absorbing the general shape of the form, owing nothing to anyone in particular. The more accurate version is that it is an industrial process that copied a specific, now-documented pile of recordings first, discarded the names attached to them second, and left the question of permission to be answered years later, long after the tracks were already in your build.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
R

Rio Castellanos

Producer & Mix Engineer

Rio Castellanos tests AI music generators against real client briefs — stems, mixes, and export quality — drawing on years behind the desk in working studios. More by Rio Castellanos →