The number is 30,117.
It is not a survey result, a market estimate, or a projection from an analyst deck. It is a count of rows in an exhibit — a schedule of sound recordings attached to a copyright complaint against Udio, one of the companies that has spent the past two years serving as the test case for AI music generation. Sony Music, in the filing reported at the time of writing, listed 30,117 recordings it says it owns and says were copied without a license to build a model that produces finished-sounding songs on demand.
Thirty thousand rows is a lot of spreadsheet. It is also, and this is the part worth sitting with, far smaller than the thing it stands in for.
Here is the verdict up front: 30,117 is a litigation-shaped number. It measures what one rights-holder can document ownership of and is willing to defend on the record. It tells you very little about how much music went into the model, nothing about whether the outputs infringe, and nothing at all about what any individual artist will see. The rest of this piece unpacks those three gaps, because the gap between what a number counts and what people think it counts is where bad business decisions get made.
What a schedule of works actually is
To sue for copyright infringement in the United States, you need two unglamorous things: ownership (or exclusive control) of the work, and a registration with the Copyright Office. The schedule is where a plaintiff proves both, one row at a time. Title, artist, registration number, sometimes a release date. Lawyers assemble it by running the catalog against registration records and pulling out everything with a clean chain of title.
That process is a filter, not a census. Recordings with disputed ownership come out. Recordings licensed in from a third party come out, because the defendant will argue the plaintiff lacks standing. Anything registered late, or never registered, comes out — or at least comes out of the part of the claim where the money is. What survives is the subset a plaintiff is confident it can carry through discovery without handing the other side a defense.
This is also why named artists show up in coverage of these filings. A schedule with recognizable names attached does rhetorical work that a raw count cannot. But the names are a sample of the schedule, and the schedule is a sample of the catalog, and the catalog is a sample of what a model may have consumed.
So read 30,117 as a floor on the claim, not a ceiling on the conduct. It is the number of doors the plaintiff has chosen to knock on.
The multiplication everyone does and nobody should
The arithmetic is irresistible, so let's do it and then take it apart. Statutory damages for willful infringement top out at $150,000 per work. Multiply: 30,117 × $150,000 = $4,517,550,000.
That number will never be paid, for five structural reasons.
$150,000 is a ceiling, not a rate. The statutory range runs from $750 to $30,000 per work at the ordinary tier. The willful tier raises the top, and a finding of innocent infringement can push the floor down to $200. A jury picks a figure inside that range. Nothing obliges it to pick the maximum, and juries in music cases have picked all over the map.
Willfulness has to be proven. It is a state-of-mind question — did the defendant know, or recklessly disregard, that the conduct infringed? A defendant with a documented, contemporaneous fair-use analysis from counsel is in a different position than one without.
Registration timing gates the whole remedy. Statutory damages and attorney's fees are unavailable for infringement that began before the work was registered, subject to a grace window for recently published works. For a catalog spanning decades this rarely bites hard, but it can carve rows out of a schedule.
"Work" is a contested unit. The statute says all the parts of a compilation constitute one work for statutory damages purposes. Whether an album counts as one work or a dozen has been litigated more than once, with courts landing differently. Applied to a 30,117-row schedule, that question alone can move the exposure by an order of magnitude.
And these cases settle. Enormous statutory awards have historically been reduced on remittitur or due-process grounds, and the prospect of that reduction is priced into every negotiation. A damages demand in a prayer for relief is a pricing anchor wearing a suit.
Which is the honest read: 30,117 is less a measurement of harm than a measurement of leverage.
What the number does not measure
It does not measure the training set
Nobody outside the defendant's engineering org and, eventually, the discovery process knows what went into the model. The plaintiff's schedule is bounded by its own catalog. A generative audio model trained to cover blues, drill, bossa nova, chiptune, and library cues has consumed material from far outside any single major's holdings — indie catalogs, production-music libraries, uploads from bedroom producers, public-domain field recordings, and material whose provenance nobody tracked. The reported allegation that songs were acquired by pulling audio from video platforms, if borne out, describes an acquisition method that scales indiscriminately. Indiscriminate acquisition does not produce a corpus that looks like one label's schedule.
It does not measure the outputs
This is the fault line the whole dispute runs along. A claim about ingestion says: copying occurred when you built the training set. A claim about outputs says: the thing your model produces is substantially similar to my recording. They are different cases with different proof burdens, and a schedule of 30,117 inputs speaks to the first.
The fair-use defense lives here too, and it is weaker in music than the text-and-search precedents suggest. The transformativeness inquiry asks, after the Supreme Court's 2023 Warhol ruling, whether the new use serves substantially the same purpose as the original. A search index that returns snippets to help you find a book serves a different purpose than the book. A model that generates a 90-second instrumental bed for a client's product video serves precisely the purpose that a production-music license serves. That collision lands on the fourth factor — market harm — which is the factor courts have historically weighted most heavily. Calling the training a fair use requires an argument that survives that collision, and "we did not have permission but the result is new" is not that argument.
It does not measure the composition, or the voice
A recording is a stack of rights, and a schedule of sound recordings covers one layer of it.
| Layer | Typically controlled by | What a master-recording schedule captures |
|---|---|---|
| Sound recording (the master) | Record label | This is the layer being counted |
| Musical composition (melody, lyrics) | Publisher and songwriters | Not counted; separate claims, separate plaintiffs |
| Performer's voice and likeness | The performer, via state law | Not counted; some states have moved to address synthetic voice specifically |
Songwriters and publishers hold claims that no label schedule asserts on their behalf. Vocal cloning sits outside federal copyright entirely in most framings and lands in right-of-publicity law, which varies state by state. If you are tracking the exposure of a generative audio company, the master-recording count is one of at least three numbers, and the other two have not been filed with the same visibility.
It does not measure what reaches an artist
Money that flows to a label under a settlement or license is not money that flows to an artist. What an individual performer receives depends on the language of their recording agreement — how it treats licensing income versus sales income, whether the deal predates the category entirely, and whether an unrecouped balance sits in the way. When a settlement is announced as a licensing arrangement rather than a cash judgment, the allocation question gets harder to trace, not easier. The label side of this fight is defending artists' recordings and is also the party with a long history of accounting disputes with those same artists. Both things are true, and only one of them shows up in the press release.
How I'd read the next complaint
More of these are coming. Six things I check, in order:
Ingestion or output. Does the complaint plead copying at training time, substantially similar outputs, or both? Output claims are harder to plead and much stronger if proven, because they don't depend on winning the fair-use argument about training.
Schedule length relative to catalog. A short schedule from a large catalog signals a plaintiff protecting its chain of title. A long one signals confidence.
Registration status of the rows. Buried in the exhibit columns, and it determines which rows can carry statutory damages.
Who is on the caption. Labels only, or performers, songwriters, publishers, and session players too? Individual plaintiff or certified class? A class action reaches a different population than a major label's suit reaches.
The remedy requested. Damages are a price. An injunction, or a demand that models trained on infringing material be destroyed, is a product roadmap question. The second one changes what gets built.
Whether the resolution carries a license. A settlement that ends with a catalog deal and a co-branded product is not a loss for either party. It is a distribution agreement that started as a lawsuit.
The fracture the number sits inside
The reason this particular filing reads as a signal rather than routine docket traffic is who did not file it. Two of the three majors reportedly reached arrangements with the same company. One did not. That is a fork in institutional strategy, and both branches are defensible: licensing converts an uncertain verdict into a known revenue line and a seat at the product table, while litigating preserves leverage on the terms of every deal that comes after.
Plaintiffs will point out that signing a license in year two does not retroactively authorize what happened in year one, and that argument is not rhetorical — it goes to willfulness. Defendants will point out that the market cleared a price, and that a functioning license market undercuts the claim of irreparable harm.
For anyone building on these tools, the practical read is simpler: the market is splitting into models with catalog agreements behind them and models racing to acquire those agreements. That distinction will show up in your terms of service before it shows up in a court opinion.
What this changes at your desk this week
If you have a cue due Friday, the litigation is background noise and the paperwork is not. Concrete checks:
- Read the commercial-use grant on your actual tier, not the marketing page. The question that matters is whether the license survives cancellation. Some tools tie usage rights to an active subscription, which means a cue you delivered in March can become a problem when you downgrade in September.
- Look for an indemnification clause and find its cap. An indemnity capped at twelve months of fees is a gesture, not a shield. Its presence still tells you how the vendor prices its own risk.
- Take the stems and the 48kHz WAV, every time. If a track has to be replaced later, stems make surgical replacement possible instead of a full re-score. A lossy render tied to an account is the worst delivery format for exactly this reason.
- Keep a per-asset paper trail: the prompt text, the date, the tool and model version, and a saved copy of the terms in force that day. When a client's clearance department asks in three years what produced the 38-second bumper, "an AI tool" is not an answer that closes the ticket.
- Tier your risk by delivery. A loop for a personal stream, a game shipping to consoles, and a national broadcast spot carry different exposure. The last one goes through people whose entire job is asking where the music came from.
Who should be watching, and who can skip it
Watch closely if you own or administer a catalog, if you clear music into ads, games, film, or television, if you build on these APIs, or if your recordings sit inside a major's schedule. Also watch if you sell production music, because the fourth-factor argument is being made about your market on your behalf.
You can skip it if you are generating loops for personal projects, or if you work exclusively inside a tool with a licensed catalog and a written indemnity you have actually read. And skip the daily coverage if you are waiting for a settled answer — appeals in copyright cases run for years, and the interim rulings that generate headlines are usually procedural.
What this piece did not answer
I cannot tell you what is in any model's training corpus; that is discovery material and most of it will be filed under seal. I cannot tell you the terms of the deals the other two majors made, because those are confidential and the announcements named partners rather than mechanisms. I do not know how a jury weighs market harm when the substitute product is a plausible one, and neither does anyone until one of these reaches a verdict rather than a term sheet. And I have no visibility into what independent labels, session musicians, and songwriters recover from any of it, because none of them are on these captions.
Where to look next: the exhibits themselves, which are public on the dockets and free to read through RECAP and CourtListener — the schedule is more informative than any article about the schedule. Then the summary-judgment briefing on fair use, which is where the real argument gets made rather than asserted. Then the version history of the terms of service on whatever tool you actually use, which changes quietly and matters immediately. And when the next deal is announced, look past the headline figure for whether it names a payout mechanism, because an amount without a mechanism is a number that stops at the label.
The figure that governs your Friday deliverable is not 30,117 — it is whichever line of your tool's license is in force on the day you export the file.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.