Home/ Articles/ The 20 kHz Shelf: What a Spectrogram Actually Proves About AI Music Generation
Production

The 20 kHz Shelf: What a Spectrogram Actually Proves About AI Music Generation

The screenshot landed in a group chat at 1am with three words attached: look at that. It was a spectrogram of a track that had done well enough for strangers to have opinions about it, and someone had…

A close-up photograph of a high-resolution computer monitor in a dark studio control room…

The screenshot landed in a group chat at 1am with three words attached: look at that.

It was a spectrogram of a track that had done well enough for strangers to have opinions about it, and someone had drawn a red circle around the top of the image. At 20 kHz the picture stopped. Not faded, not rolled off — stopped, a ruler-straight horizontal line with black above it and a dense, perfectly normal-looking mix below. To the producer who sent it, that line was proof of AI music generation. To me it was proof of something considerably more boring.

Here is the verdict, since this is the sentence that gets screenshotted: a hard shelf at 20 kHz is solid evidence that audio passed through something with a bandwidth ceiling, and close to zero evidence about who wrote the song. The number is real. The inference stacked on top of it is not.

That gap is worth sitting inside for a while, because it is the whole authenticity argument in miniature. Both sides of this fight have picked a number to stand behind, and neither number measures the thing they actually care about.

The line everyone is pointing at

If you have not spent much time staring at spectrograms: time runs left to right, frequency runs bottom to top, and brightness is energy. Kick drums are the fat orange smear along the floor. Hi-hats and cymbals are the fizzing cloud up top, which should thin out gradually as it approaches the ceiling of the file, the way a real cymbal's energy trails off into nothing.

A sharp horizontal edge in that top region looks wrong on sight, and the reason it looks wrong is that acoustics do not produce straight lines. Rooms do not have straight lines. Microphones do not have straight lines. A perfect edge means something mathematical happened. When you see energy right up to 20,000 Hz and then absolute silence at 20,001, you are looking at a decision made by software, and the instinct that says something processed this is correct.

The instinct that says and that something was a generative model is where the reasoning falls apart, because a dozen mundane processes leave a nearly identical mark.

What the shelf actually measures

Start with the most common one. Lossy encoders throw away the top of the spectrum on purpose — it is expensive to encode, most adults cannot hear it, and dumping it buys bits for the parts you can hear. A low-bitrate MP3 can cut somewhere in the mid-teens of kilohertz. A high-bitrate MP3 or AAC file typically lives right around 19 to 20 kHz. That is not a defect; it is the codec working exactly as designed. Any file that has been through a streaming platform's transcoder, a messaging app, a video export, or somebody's re-upload chain will carry a ceiling like this. If you are analysing audio you pulled off a streaming service or a video site, you are analysing the codec's output, not the master.

Then there is the sample rate. A 44.1 kHz file cannot represent anything above 22.05 kHz at all, and material is deliberately filtered below that during conversion. Anything sample-rate converted from a higher rate down to 44.1 arrives with a filter edge baked in, put there by the converter, whose steepness depends on which converter and which quality setting. Different tools, different edges, same visual signature.

And yes — many generative audio systems do render through a neural codec or vocoder that reconstructs a waveform from a compressed internal representation, and those reconstructions frequently carry a bandwidth ceiling of their own, along with a certain smeared quality in the high end that trained ears do pick up on. That is a real property of how a lot of these systems work as of writing. The specifics vary by tool and change with every release, which is exactly why you should not treat any particular cutoff frequency as a fingerprint for any particular product.

So the honest reading of the evidence looks like this:

What people read into the shelf What it can actually support
"This was machine-generated" "This audio passed through something bandwidth-limited — a codec, a converter, or a vocoder"
"The vocal is synthetic" Nothing vocal-specific; the shelf sits across the entire mix, drums included
"The artist didn't write this" Nothing whatsoever
"It's low quality" Nothing audible to most listeners; the cut sits above the top of adult hearing

There is one more problem with the shelf as evidence, and it is fatal. It is trivially removable. A pass through any harmonic exciter, a touch of dither, a tape emulation, a re-render at a different rate — the edge softens or disappears, and none of that requires intent to deceive. It is what mastering does. Which means the spectral tell only catches people who did not touch the file after export. As forensics go, that is a test that finds the careless, not the dishonest.

A photograph of a sound engineer seated in a dimly lit mixing suite, seen…

The defense always arrives as a percentage

Watch enough of these arguments and you notice the shape of the response. An artist gets accused, and somewhere in the reply there is a number. AI was maybe twenty percent of it. It was one tool out of thirty. I used it for the demo and then re-cut everything.

The percentage is doing an enormous amount of work and it is never defined. Twenty percent of what?

Twenty percent of the runtime — meaning a generated section sits in the bridge and the rest is played? Twenty percent of the track count in the session — meaning eleven of fifty-five channels came out of a model, but those eleven are the lead vocal, the topline synth and the drum bus? Twenty percent of the time spent, which counts the hours of comping and tuning after the fact but not the thirty seconds where the actual composition arrived? Or twenty percent of the decisions, which is the only version of the number that would tell you anything, and the one nobody has a method for computing?

The percentage frames the question as volume when the question is position. A generated pad under a chorus you wrote is not the same object as a generated chorus you produced a pad for, and both of them can be called twenty percent with a straight face. When a producer says the model only did a fifth of the work, the thing worth asking is not how much but which fifth, and specifically whether the melody, the lyric and the vocal performance were on the human side of the line.

I want to be fair here, because I use these tools. A lot of generated material genuinely is scaffolding. I have used a model to rough out an eight-bar bed at 92 BPM so a director could sign off on a cue direction before I spent a day playing it properly, and nothing about that arrangement troubles me. Prompt-roulette is real — you burn thirty renders to find one that is not mushy in the low mids, and the one you keep still needs surgery. The work after the render is real work. But the labour that comes after does not retroactively make you the author of what came before, and that is the sleight of hand the percentage performs.

The AutoTune comparison, held up to the light

The other standard defense is the analogy: pitch correction was called cheating too, sampling was called cheating too, the drum machine was going to put drummers out of work. All true, and all worth remembering — the history of this industry is a history of tools being denounced and then absorbed, usually by the same people.

But the analogy holds only at the level where every tool looks the same, which is the level where you stop looking at what the tool is asked to do.

Pitch correction operates on a performance. Someone stood at a microphone, made choices about phrasing and breath and where to sit behind the beat, and the software moved some of those choices onto a grid. Hard-tuned to hell, it is still a transformation of an event that occurred. Sampling is similar in structure: the source existed, someone chose it, and the choosing is the craft — the reason a great sample flip is a great sample flip is the ear that found the four bars.

Generative audio starts a step earlier. There is no event to transform. The prompt is a request, and what comes back is a performance that did not happen, by a singer who does not exist, playing a part nobody wrote. You can absolutely make art with that — the selection and editing and arrangement of generated material is a real practice with real taste in it, and I would rather engage with it honestly than pretend otherwise. But describing it as the same category of intervention as a tuning plugin requires you to ignore what each one is being handed and what each one hands back.

Which is why these public demonstrations keep going sideways. An artist opens the session to prove there is nothing to see, the video shows the input and the output side by side, and the distance between them is the entire argument the critics were making. The demo becomes the exhibit.

What the number cannot see

Even a perfect detector — assume for a moment somebody builds one that survives mastering — would leave every question people actually care about untouched.

It would not tell you about consent. The disagreement over what these models were trained on, and whether the people who made that material agreed to it or were paid, is a legal and ethical question that no amount of spectral analysis addresses. A detector tells you a model was used. It cannot tell you what the model was built from.

A photorealistic overhead photograph of a professional recording studio desk at night, an open…

It would not tell you about disclosure, which is the part that actually determines whether anyone was misled. Platforms and rights bodies have been moving toward asking for AI involvement to be declared at upload, and the specifics of those policies are moving targets worth checking rather than quoting. But the harm people are angry about is not that a machine was involved; it is that a machine was involved and the marketing said otherwise.

It would not tell you about labour — whether a session player got a call, whether a studio got booked, whether the credit list on the back of the record has real people on it. That is the material stake for most working musicians, and it is invisible in the waveform.

And it would not tell you whether the record is good. This deserves saying plainly, because the authenticity fight keeps trying to collapse the two. Plenty of fully human records are boring. Some generated material is genuinely striking, particularly in texture and sound design, where the models' tendency toward strangeness is a feature. The provenance question and the quality question are separate questions, and answering one has never answered the other.

How I'd actually decide

If you have to form a view on a specific record — you are a journalist with a deadline, an A&R with a signing decision, a competition judge — here is how I would rank the available evidence, strongest first.

Project files and stems. The session is the only thing approaching proof. Not a bounce, not a stem pack exported after the fact, but the project with its edit history, takes, comp lanes and plugin states. Someone who tracked a vocal has forty takes and a comp. Someone who generated one has a stereo file and a lot of processing on it. This is decisive and almost never available to anyone outside the room, which is precisely why the argument stays unresolved in public.

Revision on demand. Ask for a change that requires authorship: the second verse a fourth up, the same topline over a different progression, a live pass of the hook. A writer can do it, slowly. A prompt cannot be asked for it with any precision, and the failure mode is obvious to anyone in the room. This is the closest thing to a live test that exists.

Vocal micro-timing and breath. Real singers breathe in inconvenient places, drift behind the beat when the line is long, and get quieter at the end of a phrase because they are running out of air. Generated vocals have been getting better at this fast, and consonant behaviour and sibilance are where the seams tend to show as of writing. Treat this as suggestive, never conclusive, and expect it to weaken every year.

Timeline plausibility. Twelve finished, fully-produced masters in a month from someone who took a year over their last four is a reason to ask questions. It is also a reason to ask whether they had a co-producer and a deadline. Context, not evidence.

Spectral forensics. Where our 20 kHz line lands, near the bottom, for every reason above: it catches codecs more often than models, and it vanishes the moment anyone masters the file. Useful for generating a question. Useless as an answer.

Declared disclosure. Self-reported and therefore gameable, and still the only method that scales past one record at a time. Every other item on this list requires access nobody has at volume.

Notice what that ranking implies. The strongest evidence is social and procedural — files, credits, the ability to do it again — and the weakest is technical. We keep reaching for the spectrogram because it feels objective and it fits in a screenshot, which is a poor reason to trust it.

Who this argument is for, and who can skip it

If you make library or sync music, this is your business, immediately. Briefs increasingly carry provenance language and the person who can document their process has an advantage over the person who cannot. Keep your sessions. Keep your credit lists honest.

If you are covering this beat, the useful move is to stop asking whether AI was used and start asking which decisions were delegated, then asking for the artefacts. The first question gets you a percentage. The second gets you a story.

And if you are a producer at home trying to finish a track by Friday: you can skip the whole thing. Nobody is coming to audit your bassline. Use what you use, say what you used when it matters, and put your name on the parts you actually made.

The myth is that you can catch a machine-made record by looking at it, because somewhere in the file there is a line that tells the truth about who made the song. The more accurate version is that the file only ever tells you what the audio passed through on its way out, and the only thing that tells you who made the song is which decisions the artist can still account for.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
N

Nova Reyes

Editor, The Signal

Nova Reyes edits The Signal and reviews AI music tools after a decade scoring indie games and short films; still owns four broken synthesizers. More by Nova Reyes →