The screenshot is always the same shape. A detector's web interface, a waveform somebody dragged in off a rip, and one number set in a font size that does not invite discussion: 97% AI-generated. It goes up under a track that dropped four days ago. A producer with a following quote-posts it. By dinner the AI music generation allegations have a hashtag, a defense squad, and a guy in the replies who has clearly never bounced a stem in his life.
I want to sit with that number for a minute, because almost nobody who shares it can tell you what it measured. Not the accusers, not the accused, and — this is the part that should bother you — not always the person selling the detector.
Where the number comes from
Audio detectors mostly work the same way under the hood. Someone assembles two piles of files: renders from known generative models, and recordings believed to be human-made. A classifier gets trained to tell the piles apart, usually on spectrogram slices rather than raw waveform, because the artifacts people chase — smeared transients, cotton-wool high end above 14k, reverb tails that decay too smoothly, a stereo image that collapses toward mono when you flip the phase — show up more legibly as a picture than as a wiggle.
Feed the trained model a new file, it returns a confidence score. That score is the 97.
Vendors publish accuracy figures for these things, and those figures are usually real. They are also measured on held-out slices of the vendor's own data: the same model families, the same genres, the same bitrates, often the same era of generator. That's a legitimate result and a narrow one. Nothing about it promises the same behavior on a 2019 SoundCloud bounce that got mastered by a $30 online service.
What an AI detection score actually measures
It measures resemblance, not origin. A detector cannot observe how a track was made. It sees a spectrogram and reports how deeply that spectrogram sits inside the region its training data labeled machine-made. "97% AI" means, at best: this file's texture looks like the files I was told were synthetic. It does not mean there is a 97 percent chance a machine made it. Those are two different sentences, and the whole discourse lives in the gap between them.
The cleanest way to feel the gap is arithmetic. Say a detector wrongly flags 2 out of every 100 human tracks, and say 5 percent of what it's pointed at is genuinely AI. Point it at 10,000 songs: it flags roughly 190 innocent humans and roughly 500 actual renders. A given flagged track is then real about 70 percent of the time, not 97 — and that's with generous made-up inputs. Those numbers are illustrative, not measurements; plug in your own and the shape holds. As the honest-human pile gets bigger relative to the synthetic pile, the flags get less trustworthy, no matter how confident the interface looks.
What the number doesn't measure
Who did what. The interesting releases are not fully synthetic. They're hybrids: an AI-generated pad buried three layers down, a human topline over generated drums, a real bassist tracking against a machine sketch, a vocal comp run through a model to fix two syllables. A single confidence score has no vocabulary for any of that. It returns one number for a collaboration.
Production choices that look synthetic. Every trait detectors lean on is also something a human can do on purpose or by budget. Hard-quantized MIDI. Preset-y stock instruments. A limiter slammed to −6 LUFS integrated with the transients pulverized. Loops from a sample pack that a thousand other producers also bought. Everything printed dry at 44.1 because the client wanted a WAV yesterday. Cheap AI mastering applied after a human session — which is now the single most common way I've seen a clean track pick up a bad score.
Which generation it was trained on. Detectors are trained on the artifacts of the models available when the training set was frozen. Newer renders don't necessarily carry those artifacts. A detector confidently clearing a track can mean the track is human or that the generator is simply newer than the detector.
| The number is read as | What it can support |
|---|---|
| "This song was made by AI" | "This audio resembles my synthetic training examples" |
| "97% probability of AI" | "97% classifier confidence, before any base rate" |
| "The artist lied" | Nothing about intent, credits, or contracts |
| "No humans involved" | Nothing about hybrid workflows or which stem is which |
If you're the one getting accused
Detection is the wrong battlefield. You cannot prove a negative against a black box, and re-uploading to three more detectors to collect a better score is how you end up arguing about instruments instead of about your record. Provenance beats detection, and provenance is something you build before you need it.
Keep, from every session:
- The project file itself, with its edit history intact — DAW session timestamps and undo history are dull and very hard to fake retroactively.
- Dated multitrack stems at 48kHz, not only the final bounce. A mix that can be pulled apart is a mix somebody made.
- Raw takes with the room in them. Chair creak, breath before the phrase, a bad pass you kept. Nothing establishes a human faster than a mistake nobody would have generated.
- Your generation logs, if you used generative tools at all. Most serious platforms — City of Punk among them — hand you a per-track record of prompt, model, and time. Export it and keep it with the session, especially for the parts you did generate.
- A split sheet naming every contributor and what they contributed, signed at the session and not in the middle of a news cycle.
None of this is a legal opinion, and none of it obliges anyone to believe you. It changes the argument from "a website says 97" to "here is the multitrack, here is the timeline, here is who played what," which is an argument you can actually win.
The fight is about labeling, not detection
The honest read on this whole genre of story is that we're using a technical instrument to answer a disclosure question. Nobody is really asking whether a spectrogram is unusual. They're asking whether the credits are true — whether the name on the release corresponds to the work behind it, and whether anyone was going to mention the machine.
That's a metadata problem, and it's slowly becoming a platform problem. Streaming services have begun tagging fully AI-generated uploads and adjusting how those tracks surface in recommendations; the specifics differ by platform and are still moving, so check the current policy rather than anything you read six months ago, including this. But the direction is clear enough: disclosure fields, credit strings, and rights metadata are where the answer will eventually live, if it lives anywhere. A classifier score is a rumor with a progress bar.
What this piece did not answer
It did not tell you whether any specific record was made by a machine. I don't know, and neither does the screenshot, and if a detector is your only source then neither do you. It also didn't establish what share of what you streamed this week came out of a model — that number exists somewhere, keeps getting revised upward, and is measured differently by everyone who reports it.
Where to look next: the credits, not the spectrogram. Watch whether platform disclosure fields become mandatory or stay decorative. Watch whether labels start shipping provenance data alongside masters, the way film has quietly started attaching capture metadata. Watch the split sheets that surface in lawsuits, because those are the documents where somebody had to write down who actually made the thing.
A detector's confidence score is not evidence of anything except that a model has an opinion — and models have opinions about everything, which is precisely how we got here.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.