The file is 2:47 long, 124 BPM, F minor, delivered as a 48kHz WAV with stems and a master sitting somewhere around -9 LUFS integrated. It is moving in three territories this week, and somebody in your building has to decide whether it counts. That decision is what AI music generation governance exists to make, and it almost always gets framed as a question about the audio itself: is this thing synthetic or is it not. The framing feels obvious. It also has a history, and the history is shorter and thinner than the rules now resting on it.
The takedown everyone remembers, and the reason it actually worked
In the spring of 2023, a track carrying uncannily convincing soundalike vocals of two very famous artists ran up an enormous play count before it came off the platforms. Every compliance conversation since has started somewhere near that week. It is the origin story for the belief that the industry's core problem is telling machine output from human performance.
But the removal did not happen because a system listened to the file and returned a verdict. It happened the way removals have always happened: a rightsholder asserted a claim over something identifiable inside the recording, and platform terms did the rest. Reporting at the time pointed at a cleared producer tag buried in the arrangement as the lever. Whatever the precise mechanism, the enforceable object was a document — an ownership claim over a specific identifiable element — not an acoustic judgment about the whole.
The lesson the field absorbed was not the lesson on offer. What got absorbed was: we need to be able to hear these. What actually occurred was: we took it down because we could point at a right. Nearly every governance framework written since has been built on the first reading.
What the detectors were actually tested on
The second source is academic and vendor research, and it is more careful than the way it gets quoted. An audio classifier is trained on a corpus of generator outputs against a corpus of human recordings, then scored on a held-out slice of the same construction. The accuracy figures that circulate in policy decks were earned that way. A test set is not a catalog.
Here is what the classifier is often learning. Generators reconstruct waveforms through a decoder, and decoders leave fingerprints — a particular smear in the 8–16kHz band, a stereo field that stays suspiciously static, transients with a specific kind of softness on the attack. Those are real artifacts. They are also the first things to disappear when the file enters an actual release pipeline. The track gets comped against live takes, run through a limiter, resampled from 48kHz to 44.1, printed through hardware for the saturation, then transcoded to a lossy codec at delivery. Every one of those steps rewrites exactly the fine structure the model was reading.
Watermarking is the sturdier cousin and deserves credit for it. Several generators embed inaudible markers in their output, and the good implementations survive a fair amount of processing. Two limits are structural rather than technical. A watermark only marks tracks from the tools that chose to participate, and the absence of a watermark tells you nothing at all — not that the track is human, only that it did not come from a cooperating vendor, or that it was laundered through a resample. A signal that means something when present and nothing when absent cannot carry an eligibility rule on its own.
Can a detector prove a track was AI-generated?
Not to the standard a chart eligibility appeal requires. A detector returns a probability that a piece of audio resembles the outputs it was trained on — which is a statement about similarity to a reference set, not a finding about how the recording was made. That distinction survives contact with a lawyer, and probability output does not. If your rule says a track is ineligible when a classifier scores it above a threshold, your appeals process now has to defend the threshold, the training corpus, the false-positive rate on the appellant's genre, and what the score means for a record that was 80 percent tracked in a room.
That last case is the one that breaks binary policy fastest. Consider a working session: a producer records a bass performance, generates a pad bed and two textural risers from a prompt, resamples the risers into a sampler, plays them by hand, then hires a vocalist. Which category is that file? There is no percentage field in a WAV, and no honest classifier will give you one. Hybrid production is not an edge case at the margins — for game audio and post work it is already the median session.
The checkbox that became evidence
The third source of the belief is the quietest. Distributors and platforms added an AI-disclosure field to their ingestion forms. It is self-attested, typically unverified, and it was introduced as a courtesy toward transparency. Within a couple of release cycles, downstream systems began treating that field as a fact about the recording.
A self-attested flag is genuinely useful, but only at its true weight. It has the same evidentiary status as a publishing split at delivery: a claim by an interested party, actionable until contradicted, and grounds for consequences if it turns out to be false. Treated that way — a claim with a warranty behind it — the checkbox does real work. Treated as ground truth, it becomes a hole your fraud teams can drive through, because the operators most motivated to misrepresent are the ones filling in the form.
What actually survives an appeal
Run the eligibility criteria the majors and IFPI have been circulating through the question what evidence backs this, and a pattern shows up immediately. The criteria that hold are documentary. The ones that wobble are acoustic.
| The claim | What teams try to check | What stands up in an appeal |
|---|---|---|
| Training inputs were authorized | Classifier verdict on the master | The vendor's training-data warranty and indemnity in the license the creator accepted, with an effective date |
| A human authored the work | Whether it "sounds human" | Timestamped project files, stems, alternate takes, edit history, prompt and render logs |
| The AI contribution was disclosed | Metadata flag alone | Flag plus a signed creator attestation plus the tool's export receipt |
| Streams are genuine | Chart-position anomaly | Account-level telemetry, play-pattern analysis, payment and device origin |
Notice the last row. Fraud detection is the one place where forensics earns its reputation, and the reason is that it analyzes plays rather than audio — behavior leaves far more durable traces than a waveform does. That row is also where most of the actual harm sits. A synthetic track nobody streams is a curiosity; a synthetic track with a bot farm behind it is theft from a payout pool.
A workable ingestion requirement fits on one card. At delivery, require: the name and version of every generative tool used; a copy or link to the license terms the creator accepted, dated; a signed attestation of human contribution with a plain-language description of what the human did; retention of session artifacts for a defined window, producible on request; and an acknowledgment that a false attestation is grounds for removal and clawback. None of that asks anyone to hear anything.
Write rules whose evidence exists
The useful reframing is to stop treating AI-assisted and AI-generated as a property to be detected in a file and start treating them as a declared tier with escalating documentation. Low-touch use gets a flag. Heavy generative involvement gets a flag plus artifacts. Full synthesis with a licensed voice model gets flag, artifacts, and the voice license itself. The burden scales with the claim, and every tier is auditable on paper, which means every tier is defensible when someone appeals.
This also puts the pressure where it belongs: on vendors. A tool that can export a signed receipt naming its model version, the license in force at render time, and the account that generated the material is doing more for a compliance file than any detector will. We track license terms tool by tool at City of Punk because they change and the changes matter, but any team can pull the same information from a terms page and a support email in an afternoon. Ask for it before you ask for a classifier budget.
The honest version of this position is not that AI-generated music raises no problems. Unauthorized training on catalogs is a real harm with real claimants. Fraudulent streaming is a real harm with a measurable payout. Both are addressable. Neither is addressable by listening harder.
The myth is that chart eligibility for AI music is waiting on a detector good enough to sort a catalog by ear. The more accurate version is that eligibility was never an audio question — it is a chain-of-custody question, and the custody documents already exist, sitting unread in license agreements and session folders while everyone stares at a spectrogram.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.