The click was at the loop seam, and it wasn't a click.
Ninety seconds of bed music for a stealth level: 92 BPM, F minor, a detuned Rhodes over a sub meant to sit under dialogue without fighting it. The render came back clean. I dropped it into the engine, set it looping, and every eight bars something coughed. Not a digital pop — not the zero-crossing problem you fix with a two-millisecond fade. The room changed. The reverb tail arriving at the end of the loop belonged to a slightly larger space than the one the loop had started in, and my ear caught the mismatch even though the waveform ran continuous through the splice.
That cough taught me more about the tool than the nine renders that sounded fine.
So, plainly: in AI music generation, the artifacts are the most valuable thing you get. Not the hits. The tells — the smeared transient, the frequency shelf, the reverb that changes rooms, the hi-hat that dissolves into white noise on its third repeat. They are the only part of the output that cannot be marketed at you, and if you are deciding whether to pilot one of these tools this quarter, they are the fastest honest read available.
The rest of this is me earning that sentence.
Why the good renders tell you almost nothing
Every platform in this space leads with its best output. That is not dishonest, it is how product pages work, and the best output from the current generation of text-to-music models is genuinely good — good enough that I have shipped it, credited it, and not lost sleep.
The problem is that a good render is unfalsifiable as evidence. You cannot tell, from a track that sounds right, whether it sounded right because the model is strong across the space you need or because you happened to land in the middle of its training distribution. Ask for lo-fi hip-hop at 78 BPM and nearly everything on the market will hand you something usable, because that is the most densely-represented corner of the internet's music. Ask for a 7/8 taiko pattern that has to duck under dialogue, or a two-bar sting in the same key as your UI sounds, and the same model may hand you mush. The demo reel does not predict this. The artifacts do.
Artifacts are reproducible. They are properties of the architecture, the training data, and the delivery pipeline, and they persist across prompts, across sessions, and often across model versions. Once you learn to hear a platform's fingerprint, you can predict where it will fail on the day you have a deadline — which is the actual question a pilot is supposed to answer, and the one that gets buried when a trial turns into everyone generating fun songs about the CEO for a week.
There is a second reason to care, and it matters more to the people signing the invoice than to the people in the DAW. Artifacts leak information about the business. A hard shelf at the top of the spectrum tells you the model runs at a reduced sample rate, which tells you something about what a render costs the vendor to produce, which tells you something about how durable the current pricing is. Stems that bleed into each other tell you the tool is separating a finished stereo mix rather than composing in parts, which tells you adaptive and interactive audio is not on the near roadmap regardless of what the roadmap says. You are reading the technology's constraints instead of its intentions.
How I listen, and how I'd run a pilot
My method is unglamorous and takes about twenty minutes per platform. Named criteria, in the order I apply them:
Same prompt, five renders. Not five prompts. The same prompt, five times, because variance is the thing you are measuring. A tool that produces one excellent take and four unusable ones has a different cost profile than a tool that produces five solid-but-plain takes, and only one of those profiles survives contact with a client revision cycle.
Delivery format, checked and not assumed. What actually lands on disk — MP3, WAV, sample rate, bit depth — and whether the export tier you would really be on gives you the lossless file. This changes by plan and by month; check the plan you are on at the moment you are on it rather than trusting a comparison table, mine included.
A spectrogram before an opinion. I open every render in a spectrum view before I let myself have a feeling about it. Details below.
Mono fold-down and a correlation meter. Because a track that impresses in headphones and collapses on a phone speaker is not a track, it is a demo.
The loop test, even for linear work. Butt the render against itself and listen at the seam. Loop behaviour exposes internal consistency faster than anything else I know.
Stems, soloed. If the platform offers them, solo each one and listen for ghosts of the others.
Twelve-month cost, not monthly cost. Multiply, then ask what happens to everything you generated if you stop paying. Commercial-use terms in this category are commonly tied to plan tier and sometimes to plan tier at time of generation — that clause is where creators get burned, and it is worth reading before it is worth arguing about.
For a product or marketing team piloting this: give three people the same brief, the same prompt, and a shared folder, and require that every candidate track ship with its spectrogram screenshot and its mono check. It sounds bureaucratic. It takes ten minutes per track and it converts "the team liked it" into something you can defend when legal or the audio lead asks.
The artifact taxonomy
These are the tells I hear most often, what causes them, and what each one costs you downstream.
The seam
Generative audio models produce a continuous stretch of sound, not a musical structure with a defined beginning and end. That means the last bar and the first bar were never designed to meet. Butt them together and you hear it: the reverb tail is wrong, the noise floor steps, the stereo width changes by a few degrees, a hi-hat lands a handful of milliseconds ahead of the grid because tempo drifted a fraction of a BPM across ninety seconds.
The drift is the part people miss. Renders are usually close to the requested tempo, not locked to it. Over eight bars it is inaudible. Over a two-minute loop in a game build it accumulates into a click track that no longer agrees with your footstep layer.
Cost: if you need looping audio, budget editing time per asset. This is the single largest hidden labour cost in using these tools for game and installation work, and nobody's pricing page mentions it.
The shelf
Open a render in a spectrogram and look at the top of the picture. If there is a hard horizontal line — a clean edge above which nothing exists, and above which the noise also doesn't exist — the audio has been upsampled. The model generated at a lower rate and the delivery pipeline resampled to 44.1 or 48 kHz on the way out. The file's header says one thing; the picture says another.
How this sounds, if you want to hear it rather than see it: cymbals lose their air and sit slightly behind where you expect them; acoustic guitar loses fingernail; the whole mix reads as a hair darker and smaller than its loudness suggests. On laptop speakers, invisible. On a decent monitor chain in a quiet room, obvious once you know.
Cost: mostly aesthetic, until you try to sit the track under a voiceover recorded at full bandwidth in the same edit, at which point the mismatch becomes a texture problem rather than a frequency problem.
The smear
Transients are where these models still struggle most consistently. A snare should be a spike and then a body. What often arrives is a soft ramp into the spike — a millisecond or two of energy before the hit that should not be there, plus a rounded-off attack. Pluck sounds get the worst of it: a slap bass comes back sounding like a slap bass sample played through a fast compressor that has been running for years.
The practical version: if your prompt leans percussive, sparse, and dry — solo drums, a marimba figure, anything where the attack is the sound — you are asking for the hardest thing. Dense, reverberant, pad-driven material hides smear beautifully, which is one reason ambient and lo-fi renders punch above their weight in every demo reel in this category.
Cost: layering. I frequently end up putting a real one-shot on top of the generated hit — a sampled rimshot at 3 dB under, doing nothing but restoring the crack. Fifteen seconds of work, and it fixes most of the problem, which is worth saying because artifact-spotting can slide into artifact-fetishism if you let it.
The eight-bar personality
Structure is not a sound artifact, but it is a production-technique artifact, and it shows up as the thing clients describe as "it doesn't really go anywhere." Many models have a strong local sense of music and a weak long-range one. The first eight bars are convincing. Bars 9 through 16 are the same eight bars with a filter open. The bridge is not a bridge, it is a section where an element left and came back.
You can hear this most clearly by listening to a ninety-second render at 2x speed, which compresses the timescale enough that repetition becomes obvious. Arrangement decisions that felt like development at normal speed reveal themselves as variation on a loop.
Cost: for a fifteen-second product cut, none — this is invisible and the tools are excellent here. For anything with a narrative arc, you are arranging, not generating. Plan the labour accordingly.
Phantom consonants
Vocals are the hardest thing in this space and everyone who works with these tools knows it. The characteristic tell is not tuning, which is usually fine, and not timbre, which is often startlingly good. It is diction. Sibilance arrives in the wrong place. A consonant appears that belongs to no word. A vowel holds and slowly changes shape across a sustain in a way no singer's throat does. Lyrics you wrote come back with a syllable transplanted from somewhere else.
The honest version: generated vocals hold up in a mix, buried, doubled, with a delay throw. They hold up much less well solo, dry, and in front. If your use case is a hook that a listener will hear four hundred times, listen to it four times in a row before committing, because the phantom consonant you didn't notice on pass one becomes the only thing you can hear on pass ten.
Stems that were never separate
This is the artifact I care most about professionally, and it is the one that gets glossed over most often.
When a platform hands you stems, there are two possible things that happened. Either the model composed in parts and rendered each part, or it composed a stereo mix and ran a source-separation pass to pull it apart afterwards. Separation technology is good now — good enough that the distinction is not obvious from a casual listen. Solo the drum stem and turn it up, though, and you will hear a ghost of the vocal in the cymbal band, a smeared halo where the bass used to be, and a strange gated quality in the quiet passages where the separator ran out of confidence.
How to check in thirty seconds: sum all the stems back together with faders at unity and null them against the full mix by flipping the polarity of one side. Native stems tend to null deeply — near silence. Separated stems leave a distinct residue you can hear.
Cost: real and large if you need to remix. Pushing a separated drum stem up 6 dB brings the vocal ghost with it. For adaptive game audio, where the entire premise is muting and unmuting layers independently in real time, separated stems are not fit for purpose, and no amount of roadmap language changes that until the underlying model composes in parts.
Everything arrives finished
Renders come out loud, limited, and mix-complete. That is a deliberate product decision and it is the right one for most users — it is why a first-time user's first output sounds like a record instead of like a rehearsal.
It is also an artifact, and it is the one that costs professionals the most time. You cannot un-limit a master. When your job is to place music under something — dialogue, a voiceover, gameplay — you need dynamic range you can duck into, and you have been handed a track with maybe 6 dB of it. The pumping becomes audible the moment a compressor sidechains against it. The bass, already glued to the limiter, refuses to make room.
Cost: it makes generated music better at being the foreground than the background, which is the exact inverse of what most commercial workflows need.
Where I'd be honest about my own method
Artifact-hunting has a failure mode and I have been guilty of it. A spectrogram tells you nothing about whether a piece of music is any good. I have rejected renders on a visible shelf that would have worked fine under a product video watched on a phone at 40% volume, and I have shipped tracks with obvious smear because the melodic idea was worth more than the transient. Your audience is not running a null test. The listener does not know what a correlation meter is.
The discipline is worth it anyway, because it converts an aesthetic argument into a technical one. "I don't like it" loses to a stakeholder who does. "The stems bleed, so we cannot mute the drums under the VO, so this asset cannot do the job we bought it for" wins, because it is checkable.
Artifacts migrate, they don't vanish
Here is the part where I get optimistic, and I want it to be earned rather than assumed.
Every one of these tells is a snapshot of a moving target. The models that produce them are being retrained on a timescale of months, and the specific artifacts in this piece will shift — some already have. The shelf will move up as compute gets cheaper. Native multitrack output is an obvious next competitive axis and someone will ship it well. Transient handling improves visibly with each generation because it is measurable, and measurable things get optimized.
What will not happen is that artifacts go away. Every production technology in the history of recorded music has had a fingerprint, and the fingerprint has always outlived the limitation that caused it. Early digital converters produced a brittle top end that engineers spent a decade apologizing for; MP3 encoding produced pre-echo on castanets and hi-hats that codec developers chased for years; gated reverb existed because a talkback mic on an SSL desk did something unintended. Two of those became genres. Producers went looking for the artifact on purpose, and are still buying hardware to get it.
So the useful prediction is not that generated audio will become indistinguishable. It is that the current fingerprint will become a period signature — the way a certain smeared, over-limited, harmonically-dense wash reads as 2020s the way gated snare reads as 1984. Someone is going to build a career on making that sound deliberate. Learning to hear the artifacts now is how you end up being that person rather than being surprised by them.
Who this is for, and who should skip it
Run this protocol if you are placing music under other audio, shipping loops into a build, producing at volume where a systematic failure costs you a re-render across fifty assets, or standing in front of a legal team that will ask what exactly you licensed.
Skip most of it if you are making short-form social video, scoring a montage nobody will hear twice, or exploring as a hobbyist. The tools in this category are already good enough for those jobs, the artifacts described here are mostly inaudible in those contexts, and learning to hear them will reduce your enjoyment of a hobby without improving your output. That is a real cost and I am not going to pretend otherwise.
For the pilot decision specifically: the honest answer for most teams is that these tools clear the bar for foreground, short-form, one-off music, and do not yet clear it for background, long-form, or interactive music without meaningful human labour on top. Budget the labour or scope the pilot to the first category. Piloting against the second and concluding "the technology isn't there" wastes a quarter learning something you could have learned in an afternoon with a spectrum analyzer.
The decision, stated plainly
Video edit due Friday, music in the foreground, fifteen to sixty seconds: any of the major text-to-music platforms will do this, and your choice should hinge on which plan tier grants the commercial rights you need and what the file format actually is, not on output quality. That fight is close enough to call a draw as of writing.
Adaptive loops for a game build: no current general-purpose generator is a clean fit, because of the seam and the separated stems. Generate source material, then loop it, edit it, and layer it yourself — or work with a tool that composes natively in parts, and verify that claim with a null test rather than a feature list.
Podcast intro: generation wins outright, on the condition that you re-render or ride the last four seconds so the music resolves under your first line instead of fighting it. The over-limiting is your problem here more than the model's.
Try this, this week
Take the three renders you liked best — the ones you already used, or nearly used. Drop them into a free spectrum viewer (Spek works, so does your DAW's analyzer) and look at the top of the picture for a flat horizontal edge. Then sum each one to mono and listen to what happens to the low end and the reverb. Two minutes per track, no purchase, no subscription.
You will end up with one of two outcomes, and both are worth having: either your tool is delivering the bandwidth it claims and translating to a phone speaker, or it isn't and you now know exactly which stage of your process has to compensate. Either way you have stopped taking the vendor's word for what is in the file.
Look at the picture before you fall in love with the sound.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.