You have watched a version of this clip. A gear brand posts a sixty-second promo: hands on a fretboard, a warm strummed figure in open D, a bass sitting a little too politely underneath, a hi-hat that never once drifts. Around the fifteen-second mark, under the reverb tail, there is a thin fizz in the top end that does not belong to the room. The comments find it before the algorithm does. By the next morning there is a video essay with a spectrogram on screen, a red circle drawn around a smear near 12kHz, and a thread in which the phrase AI in music production is carrying an enormous amount of weight for a group of people who have not heard the session files.
The verdict, stated plainly: you cannot reliably identify machine-generated audio by ear or by spectrogram, most of the artifacts people cite as proof have mundane production causes, and the only durable signal of a brand's honesty is a disclosure policy that existed before anyone accused them of anything. Everything below is about how to hold that position without becoming either a mark or a member of a mob.
I have skin in this. I score indie games and short films, I own four synthesizers in various states of disrepair, and my delivery folders routinely contain stems that were separated by a model, a vocal that was pitch-corrected, and a first-pass master that a machine roughed in before I took over. If the standard is "no algorithm touched this," I fail it. So does nearly every record you love from the last fifteen years.
What most people do
Most people run a test they believe is forensic and which is, in practice, vibes with a waveform attached.
The method as it is actually practiced in comment sections and teardown videos goes like this. Rip the audio off the video. Load it into an editor. Open the spectrogram. Hunt for tells. The tell list is remarkably consistent across the community, which is part of why it feels authoritative:
- A noise floor that misbehaves. Hiss that appears and disappears with the arrangement rather than running underneath it, or a bed of fine hash in the 10–16kHz region that has no obvious source.
- Transients that smear. A pick attack that arrives soft-edged, a snare whose first millisecond looks blurred rather than vertical.
- A grid that never breathes. Hi-hats landing on the same subdivision to the sample for ninety seconds, no push or drag anywhere in the pocket.
- Stereo that collapses oddly. A wide, gauzy image that goes strange in mono, with phase artifacts nobody would print on purpose.
- Hands that do not match the audio. The player's fingers move to a different voicing than the one you hear, or a chord change lands a frame or two off.
Stack three or four of those and a listener feels certain. The certainty is the problem. Nobody in that thread has the multitrack, the plugin chain, the session tempo, or the codec history of the file they are analyzing. They have a YouTube re-encode and a strong prior.
And here is the part I refuse to be snide about: the strong prior is earned. The people running these teardowns are the same people who bought a stock-music subscription for a client project, exported the cue, delivered the video, and then learned that "royalty-free" covered the download but not the broadcast, or that the commercial tier they were on excluded the exact use case they needed, in a footnote three clicks deep. They are the people who had a track pulled after a library changed its terms retroactively. When an industry has spent two decades burying the terms that matter, the audience does not owe it the benefit of the doubt.
So the cycle runs, and it runs the same way every time. Accusation lands. The brand issues a short denial that reads like it went through legal. The community reads that as evasion and escalates, because a short denial contains nothing falsifiable. The brand then comes back with detail — session tempo, sample rate, how many compressor instances were on the bus, who played what, why there is extra noise in the top end — and concedes some smaller point, often that an assistive tool was in the mastering chain after all. That is roughly the shape of the string-brand controversy that has been circulating as of writing, in which the company attributed the disputed artifacts to modern production practice while acknowledging a machine-assisted step it had not previously flagged. Some readers came away satisfied. Plenty did not, and their reasons were specific: the timing questions the detail did not address, the fact that the detail arrived second rather than first.
That pattern is the interesting thing, more than any individual case. Denial loses. Specificity partially wins. Doubt survives both.
What the evidence suggests
Every tell on the list has a boring explanation
This is where the folk-forensic method falls apart. Not because the artifacts are imaginary — people are hearing something real — but because each one has a well-documented, decades-old cause that has nothing to do with generative audio.
| The tell | What people take it to prove | What it actually indicates |
|---|---|---|
| Hash in the top octave | Model output artifacts | Aggressive limiting, saturation stages, sample-rate conversion, or the platform's own lossy encode |
| Noise floor that pumps | Synthetic "room" pasted in | A gate or a fast bus compressor riding a stack of layered stems |
| Smeared transients | Model can't render attack | Multiband processing, clipper on the master, or a re-encode at low bitrate |
| Perfect grid | Nothing human played it | Quantized MIDI, drum replacement, or a programmed bed under a live overdub |
| Collapsed mono image | Model stereo imaging | Mid-side widening, chorus on a bus, or a reverb with decorrelated tails |
| Hands out of sync with audio | The audio was never played | The performance was re-recorded after filming, or the video editor cut to a music bed — a practice older than the internet |
None of that clears anyone. It means the artifact is evidence of processing, and processing is not the same claim.
Detectors exist, and they are not evidence
There are real classifiers for synthetic audio, and the research is legitimate. The honest caveat is the one that matters to you: published accuracy figures in this field are generated on curated test sets, where the synthetic examples come from a known list of generators and the real examples are relatively clean. Promo audio is neither. It has been compressed, limited, loudness-normalized by the platform, and re-encoded at least twice before it reaches your spectrogram.
The failure mode that should worry you is not the miss. It is the false positive: a genuine performance, heavily processed, flagged as machine-made. In a research paper that is a percentage point. Applied to a named company or a named session player, it is somebody's reputation. I have not seen a detector publish a false-positive rate on heavily-mastered commercial audio that I would be willing to stake an accusation on, and until one does, detector output belongs in the category of "reason to ask a question," not "finding."
Provenance metadata is the right idea with the wrong failure mode
The structural fix is content credentials — cryptographically signed provenance attached to the file, describing what tools touched it. The C2PA specification exists, it is backed by serious players, and it is the correct architecture. It also has a failure mode that gutted every metadata standard before it: the credential does not survive the pipeline. Re-encode the file, bounce it through a video editor, upload it to a platform that strips metadata on ingest, and the signature is gone. Absence of a credential proves nothing, which means presence of one is the only usable signal, which means the standard only works when the whole chain cooperates. Adoption is uneven and will stay uneven for a while.
"Did they use AI" is not a yes-or-no question
This is the part the outrage cycle cannot metabolize. Consider a plausible chain for a gear demo: a real player, tracked live; a comping tool that assembled the best take across passes; a de-noiser on the room mic; pitch correction on a sung line; stem separation used to rescue a reference; a mastering assistant that set the first EQ curve and ceiling. Where in that chain does the video become dishonest?
The field has argued this before and reached no clean answer. Pitch correction was the same fight in 1998, sample libraries were the same fight before that, and the drum machine was the same fight in 1980. Each time the line settled not on the tool but on the claim: what did you say you did? Nobody accuses a producer of fraud for using a sampled orchestra. They would accuse one for putting a photograph of a string section on the sleeve.
Which means the ethical question is a disclosure question, not a signal-processing question — and disclosure is a document, not a waveform.
Specificity persuades, and its limits are real
The reason the detailed corporate response works better than the denial is that specifics are falsifiable. "We recorded this at 87 BPM at 44.1kHz with a named player on a named instrument, and here is the plugin chain" gives skeptics something to check. "We stand behind our content" gives them nothing, and an audience with nothing to check assumes the worst.
But be fair in both directions: specifics can be manufactured too. A plugin count is not an audit. The thing that actually distinguishes a credible response from a well-drafted one is whether the underlying material can be produced on request — the session, the stems, the raw take. Detail that cannot be checked is rhetoric with numbers in it.
What I actually do
Here are the criteria I use, in the order I use them, when I need to decide whether to trust a brand's demo, a library's catalog, or a collaborator's delivery. This is a workflow for making decisions with money attached, not for winning arguments.
1. Check whether the credits are falsifiable. A named player, a named studio, and a named engineer are things I can verify with one message. "Produced by our creative team" is not a credit; it is the absence of one. This costs nothing and sorts most cases. When a company that has always credited its session musicians stops crediting them on one spot, that gap tells me more than any spectrogram.
2. Ask for one unprocessed file. For a library or a collaborator, I ask for a single 48kHz WAV of a raw take — a DI, or a room mic with nothing on it — alongside the finished cue. Nearly every disputed artifact lives in the mastering chain, and an unprocessed take moves the conversation from inference to material. A vendor who cannot produce one has told me something, and a vendor who produces one immediately has also told me something.
3. Never analyze the re-encode. If I am going to look at a spectrogram at all, I look at the highest-quality source available, and I compare it against a known-human reference that went through the same platform pipeline. A hi-hat that smears identically in both is a codec artifact, not a confession. This single control kills the majority of "tells" people post.
4. Read the license and the disclosure page instead of the promo. This is where the actual damage happens to working people. I want to know: does the commercial grant cover broadcast and paid social, or only the download? Does it survive a subscription lapse, or do my delivered projects go unlicensed when I stop paying? Does the vendor indemnify me if a generated cue turns out to be derivative? Are stems included at the tier I can afford, or gated above it? Those terms differ enormously between vendors, they change, and they are the difference between a cue you can ship and a cue that gets a client's video pulled.
5. Weight policy over statements. A disclosure policy published before an incident is worth more than any explanation published after one. If a brand documented its stance on generative tools last year — where it uses them, where it does not, how it labels — I extend real credit. If the policy appeared the week the accusation trended, I read it as damage control and treat it accordingly. This is the single criterion that has best predicted, in my experience, whether a company will be straight with me on the second controversy.
What I do not do
I do not post detector screenshots as accusations. I do not treat a quantized grid as proof of anything, because I have programmed thousands of them myself. And I do not pretend my own workflow is pure. My delivery notes name every assistive tool that touched a cue and the license state of every source, because when a client's lawyer asks in eighteen months, the answer needs to be a document rather than my recollection. That practice has cost me exactly nothing and saved me twice.
I will also say the uncomfortable part: some brands are using generative audio in demos and not disclosing it, the incentive to stay quiet is obvious, and my five criteria will not catch a competent liar. They catch sloppiness and they reward the honest. That is the realistic ceiling.
Who this is for, and who should skip it
This approach is for anyone whose income depends on the license chain holding up — freelance composers, game audio contractors, editors licensing beds for client work, anyone who has to sign a deliverables form. If you have to answer for a cue in a contract, criteria beat instinct every time.
Skip it if you are making music for yourself. The overhead is real, the paperwork is tedious, and if nothing you make is going to be examined by a rights holder, checking a vendor's indemnification language is time you could spend on the record. Enjoy the tools and stop reading terms of service. That is a legitimate way to live.
The question I cannot answer
Here is what is genuinely unsettled, and I do not think anyone has closed it.
We do not have good evidence about whether ordinary listeners can distinguish machine-generated music from human performance under proper blind conditions, at full-mix quality, in the genres where the models are strongest. The teardown videos are not blind — the person analyzing already believes something. The listening studies that exist tend to use short excerpts, unrepresentative material, or listeners who know what the study is about, and effects that show up in that setup have a way of evaporating under stricter controls.
Which leaves the possibility that unsettles me most: that what we are detecting is not the audio at all. That the fizz at 12kHz feels synthetic because we were told to look for it, and that the same fizz on a record we already love reads as texture, as tape, as character. If the label is doing the hearing rather than the ears, then the entire authenticity fight is not about signal processing and never was — it is about who gets credited, who gets paid, and who told you the truth before you asked.
So the question I would put to anyone certain they can hear the difference is the one nobody has run properly yet: could you still tell, if nobody had told you to listen?
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.