Home/ The Signal/ Industry/ The 20 Percent Number: What One Detection Study Actually Says About AI Music Generation Enforcement
Watermarking

The 20 Percent Number: What One Detection Study Actually Says About AI Music Generation Enforcement

Every few months a number circulates through the AI music generation compliance conversation, and it is almost always a detection-accuracy figure: a platform or a research group reports that its…

A close-up photorealistic studio photograph of a recording engineer's mixing console in a dimly…

Every few months a number circulates through the AI music generation compliance conversation, and it is almost always a detection-accuracy figure: a platform or a research group reports that its watermark survives some percentage of common audio manipulations, and the number gets quoted as if it settles the enforcement question. Sometimes it's in the high nineties. Sometimes, once you get past the press summary and into the methodology, the survival rate under aggressive processing drops toward twenty percent. Both numbers can come out of the same system. The gap between them is the entire story.

This piece is about what those figures actually measure, and — more usefully for anyone tracking how platforms respond to legal pressure — what they structurally cannot measure, no matter how good the accuracy gets.

What a watermark detection rate actually measures

A watermark accuracy figure is a measurement of one narrow thing: given an audio file, can the detector tell whether this specific system's imperceptible signal is present.

That's it. It is a claim about a file, made by the party that embedded the mark, using a detector that party controls.

The conditions of the measurement matter more than the percentage. Published watermark evaluations typically test against a manipulation battery — MP3 encoding at various bitrates, resampling, time-stretch, pitch-shift, added noise, EQ, dynamic-range compression, sometimes a re-record through a speaker and microphone. The headline number is usually detection under mild processing. The interesting number is the tail: what happens at 64 kbps, after a two-semitone pitch shift, under a layer of drums.

Three things degrade a watermark in practice, and they map exactly onto how music is used:

  • Transcoding chains. A track uploaded to one platform, re-encoded, downloaded, re-uploaded elsewhere passes through multiple lossy codecs. Each pass is a small subtraction from the signal the detector needs.
  • Layering and mixing. A generated stem placed under a live bass and a drum bus is no longer the file that was marked. It is a component of a new file, at a lower relative level, competing with uncorrelated audio.
  • Deliberate removal. Watermark-stripping is an adversarial problem, and adversarial problems have research communities. Any published scheme becomes a target for the next paper.

None of this makes watermarking useless. It makes the accuracy figure a statement about laboratory conditions, and enforcement does not happen in laboratory conditions.

What detection does not tell you about training data

Here is the part that gets lost when the number travels: output detection and training-data provenance are unrelated technical problems.

A watermark answers "did system X generate this file." Every significant copyright action against generative music platforms — the German collecting society GEMA's litigation, the major-label suits filed in the United States — turns on a different question: what recordings and compositions went into the model, under what authorization. A perfect detector on the output side produces no evidence about the input side. It cannot show whether a protected master was in the training corpus, whether a licence covered it, or whether the resulting model weights constitute an infringing derivative.

Platforms are generally clear about this distinction in their own filings, where they argue training practices as fair use or under text-and-data-mining exceptions rather than pointing at their detection stack. It is the coverage in between that collapses the two, treating a transparency announcement as though it were a response to an infringement claim. It isn't. It's a response to a different pressure entirely.

The three layers, and what each one misses

Platform compliance stacks in this space have converged on three distinct mechanisms. Analysts reading vendor announcements should be able to sort claims into the right bucket, because each layer fails differently.

Watermarking embeds an imperceptible signal at generation time. It identifies the generator. It requires the generating platform's cooperation to embed and usually its cooperation to detect, it degrades under processing, and it says nothing about any model other than the one that marked the file.

Fingerprinting — the Audible Magic and Content ID lineage — computes a perceptual hash of a reference recording and matches unknown audio against a registered database. It is mature, it survives moderate processing well, and it is the layer that actually powers takedowns at scale. Its blind spot is definitional: it matches against registered references. A generated track that resembles a protected work in style, instrumentation, and feel but shares no acoustic frame with any registered master will pass clean, because fingerprinting was built to catch copies, not resemblance.

A photorealistic wide-angle photograph of a woman in her thirties standing alone in a…

Content credentials — C2PA-style signed metadata — attach a cryptographic provenance record to the file. When present and intact, they are the richest of the three. They are also the most fragile: metadata is stripped by nearly every consumer upload pipeline, and absence of a credential proves nothing, since unmarked files are the default state of the world.

Enforcement question Layer that can answer it Its failure mode
Did platform X generate this file? Watermark Degrades under transcoding, layering, adversarial removal
Does this recording copy a registered master? Fingerprint Only catches registered references; blind to stylistic resemblance
What is this file's declared origin chain? Content credentials Metadata routinely stripped on upload; absence proves nothing
Was protected material used in training? None of them Answered by discovery and disclosure, not by signal processing

That last row is the one worth pinning to the wall. The technical stack, fully deployed and working perfectly, does not reach the question the litigation is actually about.

Why platforms ship detection anyway

If detection doesn't answer the training question, the obvious analyst reading is that it's theatre. That reading is too cheap, because detection does real work — just not the work the headline implies.

It establishes a licensing substrate. Rightsholders negotiating catalogue deals with generative platforms need a mechanism for attribution and payment. Detection infrastructure is the meter. Several publicly reported label negotiations have proceeded alongside detection and attribution commitments, and it is difficult to structure a royalty flow for machine-generated output without some way of identifying it.

It establishes a regulatory posture. The EU AI Act's transparency provisions require providers of generative systems to mark synthetic output in a machine-readable way, with obligations phasing in on a published schedule. Whatever a given platform says about its motives, shipping a watermark ahead of a compliance deadline is a legible act with a legible audience.

And it establishes a good-faith record. Platforms facing infringement claims benefit from a documented history of safeguards — prompt filters that decline named-artist requests, output screening against reference databases, provenance marking. Whether that record affects liability is a question for courts, and courts have not resolved it. It affects the negotiating table now.

What to watch instead of the accuracy percentage

For anyone tracking this sector, four disclosures carry more signal than any headline detection figure:

  • False-positive rate, stated separately. A detector that flags human-made recordings as synthetic is a liability, not a compliance asset, and the false-positive rate is the number most often absent from announcements.
  • Survivability methodology. Which manipulations were tested, at what strength, and was the battery adversarial or merely typical? "Survives MP3 compression" and "survives a determined removal attempt" are different claims.
  • Detector access. Can rightsholders and third parties run the detector, or must every query route through the platform that embedded the mark? Self-verified detection has a structural credibility ceiling.
  • Evidentiary standing. Has any detection output been offered, and accepted, as evidence in a proceeding? So far the mechanisms function mainly as inputs to platform enforcement workflows and licensing arrangements rather than as adjudicated proof.

A fifth, quieter one: whether a platform's detection commitments are contractual to partners or merely announced. The first survives a change of strategy. The second does not.

The number in context

Return to the figure. A detection rate — twenty percent, ninety-nine percent, whatever this quarter's paper reports — is a measurement of a signal-processing system under specified conditions. It is genuinely useful for what it covers: platform-side moderation, catalogue hygiene, the plumbing of an attribution deal. Read as a proxy for whether generative music platforms can be held accountable for what they trained on, it is not a weak answer. It is an answer to a different question, arriving where a harder one was asked.

The myth: watermarking accuracy tells you how well AI-generated music can be policed.

The more accurate version: watermarking accuracy tells you how reliably one company can recognize its own output under lab conditions, while every question a court is actually being asked — what went into the model, under what licence, with what authorization — sits entirely outside what any detector can see.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
R

Rio Castellanos

Producer & Mix Engineer

Rio Castellanos tests AI music generators against real client briefs — stems, mixes, and export quality — drawing on years behind the desk in working studios. More by Rio Castellanos →