A composer I work with sent me a 90-second cue last spring and asked whether I could tell. Detuned Juno pad, bass sitting a hair flat under a broken 808, room tone from what sounded like an actual room. I said generated. It was him, played into a laptop at two in the morning. I got it wrong on first listen, which is the problem with any AI music policy built on someone listening.
The verdict, up front: you cannot enforce an authorship rule by judging recordings, and the platforms that keep trying end up accusing their own artists. The rules that hold up verify people at the account level and leave the audio alone.
That conclusion is not new, but the belief it contradicts is so widespread that it's worth asking where it came from. It has a source. The source is thinner than the belief.
The belief has a birthday
For roughly two years, generated audio really was audible. In the first public wave of text-to-music systems, the tells were consistent enough to teach: cymbals that smeared into a wash instead of decaying, transients that arrived soft, a reverb tail that didn't behave like any room you've stood in, a faint seam at the bar line where the model stitched, lyrics that dissolved into vowel soup by the second chorus. Producers traded these tells the way we trade plugin settings. A lot of people learned to hear them, correctly, and then quietly assumed the skill would keep working.
Two other things pushed in the same direction. Universities already had a template for this problem — the plagiarism checker — so a detector felt like the obvious shape of a solution before anyone asked whether audio was the same kind of object as an essay. And image detection had a head start, first through visible artifacts and later through watermarking proposals, so audio inherited the assumption that the trick would port over.
It ported badly. A picture is one static array of pixels that usually reaches the platform close to how it left the model. A piece of music gets bounced, gain-staged, compressed on a bus, mastered by somebody's chain, transcoded to a lossy stream format, and sometimes replayed through a monitor into a room and recaptured. Every one of those steps is a light wash over the exact fine-grained detail a detector was looking at.
The source is thinner than the belief
Here is the part that gets skipped. A detector learns the fingerprints of the generators that existed when it was trained. Its confident performance is measured against those generators, usually on raw renders. Neither condition holds on a live platform. New models ship constantly, and almost nothing arrives raw.
Then there's hybrid work, which is now most work. A songwriter cuts a topline over a generated bed. A producer replaces a generated kick with a sample they recorded off a Rhodes case. Somebody runs stem separation on their own old multitrack, which is a model touching human audio, and then replays the bass part by hand. Ask what percentage of that record is AI and you get an argument, not a number. A detector returns a number anyway, which is the tell that it's answering a different question than the one you asked.
The scale math finishes it. Suppose a classifier is wrong on one upload in a hundred, which would be a strong result under real-world conditions. On a hundred thousand uploads that's a thousand mislabeled records, and the humans among them have to prove a negative. The only artist who can win that appeal is one who kept the session — dated project files, unbounced takes, a phone video of a pass. Which is a hint about where the enforceable evidence actually lives, and it isn't in the WAV.
The economics nobody detects their way out of
Indie artists don't lose money because a listener mistakes a generated track for theirs. They lose because of denominators.
Most streaming payouts come out of a pro-rata pool: a fixed amount of money divided by everything that got played. Editorial and algorithmic discovery slots are finite by design. Membership tiers, limited runs, presale queues, numbered supporter systems — all of the scarcity structures small artists have built over the last decade — depend on the set of participants staying countable. Recorded music has been priced near zero for a long time; that fight is over. What changes with cheap generation is the size of the denominator, and every scarcity mechanism sits on top of it.
The cost asymmetry is the design constraint most policies miss. Making a record costs a person weeks. Uploading four hundred tracks costs an operator an afternoon and some compute. Any moderation system whose cost per item exceeds the uploader's cost per item loses on arithmetic, no matter how good the classifier is. Reviewing files is a recurring cost that scales with supply. Reviewing people is a one-time cost that scales with population, and population is the thing you were trying to keep countable in the first place.
That's the whole case for account-level verification, and it's an economic case, not a moral one. Nobody has to believe machine-made music is bad art for the math to work out this way.
How I'd decide, if I ran the platform
Name what you're protecting. A royalty pool, a discovery slot, a membership number, and editorial trust are four different assets, and they don't imply the same rule. If you're protecting a pool, per-account caps and payout rules do more than any authorship policy. If you're protecting membership scarcity, identity is the entire game.
Price enforcement per person, not per file. One verification of a human, once, holds across everything they ever upload. One review of a file holds for that file.
Decide what evidence you accept, and pick things that are cheap for an honest artist and expensive at volume. A short video of a take. A project file with an edit history. Stems at 48kHz that match the master. A ten-minute call. None of these prove a recording is human. All of them make bulk operation costly, which is the actual objective.
Put the line at authorship, not at tools. Pitch correction, time alignment, stem separation, drum replacement, denoise, and mastering assistants are instruments and processors, and drawing a line at them means banning half of modern production including plenty of records made before any of this was controversial. The line worth defending is whether a person wrote it and a person performed it. State that in one sentence, publicly, and enforce that sentence.
Write the revocation rule before you need it. What happens to a membership slot when an account is removed — retired or reissued? Who hears an appeal, and on what evidence? A policy without a stated failure mode is a press release.
Be honest that verification is a barrier. It slows signups, it costs staff hours, and it excludes people who can't or won't show identification. Publish that cost instead of pretending the queue is a feature.
Who this is for, and who should skip it
If you're an independent artist: your archive is now part of your defense. Keep dated project files, keep one unglamorous phone video per session, keep stems rather than only masters. It takes a few minutes and it's the only thing that answers a false accusation.
If you run a small platform or a community where status is countable — member numbers, limited pressings, a curated roster — human approval works at your size, and it works better than any classifier you could license. Do it while the roster is small enough to review.
If you run a licensing library or a stock catalog, skip most of this. Your product is the cue, not the person behind it. Your obligations are disclosure and clean license terms, and your users care about indemnification, not provenance.
The myth: with a good enough detector, a platform can sort AI music from human music and keep its catalog honest. The more accurate version: a platform can't know what made a recording, but it can know who stands behind it — and that was what its economics were resting on the whole time.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.