Home/ Articles/ 64 Percent: What a Blind Test Really Says About AI Music Generation in Games
Adaptive Music

64 Percent: What a Blind Test Really Says About AI Music Generation in Games

Fourteen out of twenty-two. That's how many people correctly picked the human-composed loop when I dropped two eight-bar clips into a Discord full of game developers and audio people and asked which…

A professional photorealistic photograph of a dimly lit home studio at night, shot at…

Fourteen out of twenty-two.

That's how many people correctly picked the human-composed loop when I dropped two eight-bar clips into a Discord full of game developers and audio people and asked which one came out of a text prompt. Sixty-four percent. A coin flip is fifty. I'd spent the better part of a year telling anyone who would sit still that AI music generation is obvious once you know the tells, and my own evidence came back as a shrug.

I want to stay in that number for a minute, because the reflex is to explain it away. The more useful move is to work out precisely what it measured, and then everything it never touched.

What the number actually measured

Two clips, eight bars each, roughly nineteen seconds. Both 128 BPM, both in A minor, both exported as 48kHz WAV. One I wrote in a session I still have open. One came from a prompt. Twenty-two people voted; fourteen were right.

Here is what that setup can honestly support: a self-selected group of listeners, on unknown playback — I know for a fact some of them were on laptop speakers, because they told me — identified the source of a nineteen-second loop stripped of every scrap of context, slightly more often than chance. That's it. It is a poll, not a study. I made one of the two clips, which means I had a stake in the outcome and no business running the test in the first place.

And there's a bigger contamination in it. The generated clip I posted was the survivor of about thirty renders. The other twenty-nine had problems I could name: smeared transients, a stereo image that stayed pinned to the middle, a snare that sounded recorded in a stairwell, one take that drifted a quarter-tone flat across the loop point. Prompt-roulette is real and it is most of the work. So the test measured my curation at least as much as it measured the model. A person with ears sat between the generator and the audience, throwing things away. Take that person out and the number collapses.

Which is the first honest finding: the ceiling of current output is high enough to fool people in short clips, and the floor is still on the floor.

Can listeners tell when game music is AI-generated?

In short clips, mostly no — and anyone claiming the difference is obvious is describing bad output, not the category. What listeners reliably catch shows up over duration, not timbre: fills that land identically every pass, a stereo field that never opens up for a chorus, endings that fade instead of resolve, and the same four-bar idea sitting under every emotional state in the level. Detection is a function of time-on-task and context. Give someone twenty seconds and they're guessing. Give them forty hours of play and they stop guessing and start getting irritated, usually without being able to say why.

The clip is a photograph of something that's supposed to move

A game score isn't a track. It's a system that a player operates without knowing it.

The stealth cue I keep pointing at when I teach this is seven stems: a pulse, a sub, two pads, an arpeggio, a percussion tier, and a single sting layer that only exists for the moment a guard's vision cone catches you. Three intensity tiers. Transitions quantized to two bars, so when you're spotted the arp enters on the next downbeat instead of hard-cutting. The player hears maybe forty minutes of music across that level and never hears a beginning or an end. They hear state changes.

Generation hands you a rendered mixdown. Some tools export stems, and stem export is genuinely useful. Fewer will let you constrain key, tempo, and bar length across a family of related cues so the pieces are siblings rather than strangers. And there's a difference between stems that were composed to be recombined and stems separated out of a finished mix after the fact — with the latter you hear the ghost of the snare living inside the pad, and it gets loud the moment you mute a layer. Every render I've pulled apart that way has had bleed somewhere.

My poll asked whether a photograph of a system looks like a photograph of a different system. Of course it does. Both are photographs.

Revision versus replacement

Here's the failure I hit on actual paid work, every time.

The director says: the brass at 1:12 is stepping on the line read, keep the swell but pull it under. In a session, that's a fader move, six decibels, plus a shorter release on the reverb, and you're done in ninety seconds. With a prompt, you roll again. You get a different piece of music. Maybe a better one. Definitely not the same one with the brass pulled under.

Iteration in generation is replacement, not revision. That single property decides where these tools sit in a pipeline more than sound quality does. It's why generated audio has been genuinely good to me for temp tracks, greybox prototypes, ambient beds nobody will consciously hear, and pitch materials that need to exist by Thursday — and why it has never once survived to a hero cue, where notes get argued about individually.

Before an AI cue goes into a build

  • Loop it two hundred times. Not five. The seam and the repeated fill both surface around pass thirty.
  • Solo every stem. Listen for bleed from instruments that aren't in that stem. If you hear the ghost, you can't mix it.
  • Check the tail. A fade-out is not an ending. If it can't resolve on a downbeat, it can't transition.
  • Play it against the dialogue bus. Generated mids are often crowded; the render sounds full in isolation because it filled the space your voice track needs.
  • Read the license for the tier you're actually on, not the one on the marketing page.

Constrained prompts do more of the work than descriptive ones. This shape has been reliable for me:

Sparse stealth loop, 92 BPM, D minor, 16 bars, no intro,
no fade-out, resolves on the downbeat of bar 17. Detuned
analog pulse, sub-bass on beats 1 and 3, no drums, no vocals,
no melodic hook — leave headroom for a lead.

The bar count and the resolve instruction are what make it usable in middleware. The negatives do most of the lifting: no melodic hook keeps the render from fighting whatever you'll write on top of it, and no vocals has saved me more renders than any adjective I've ever typed.

The thing the poll couldn't test at all: whether you can ship it

Nobody in that Discord was voting on rights, and rights are where indie teams actually get hurt.

Terms vary by tool and by tier, and they change. The questions worth answering before a cue enters your repo: does the license cover commercial release across every storefront you're shipping to, does it survive you cancelling the subscription, does the free or trial tier grant the same rights as the paid one, and is there any indemnity if a claim lands on your soundtrack upload. Some services are explicit on all four. Some are quiet about the fourth. That silence is information.

This isn't legal advice and I'm not qualified to give any — read the terms yourself, and on a funded title, pay someone who does this for a living to read them. We track stated license terms per tool on the compare pages here, but the terms are the source and we're a snapshot; check the source.

What I'll concede, and what I won't

My poll took my favorite argument away from me. I can't tell you people hear the difference, because on that day, they barely did.

So the argument has to be the honest one. A score is a record of decisions. I have left a string sample slightly out of tune in a cue on purpose, because the character had been walking for two days and I wanted the listener tired. Nobody in a blind test would flag that. It works because someone chose it, and choices accumulate into a thing that feels authored, over hours, in a way no twenty-second sample can expose.

The counterargument I hear most is that nobody mourned when sample libraries displaced session players. I don't think that's the same trade, and I'm happy to be argued with on it: a sample library is a set of recordings that a person then has to arrange, and the arranging was always the job. What's on the table now is the arranging.

What this doesn't answer

I still don't know whether a game scored end-to-end by prompt could hold a player across forty hours — whether the irritation I described actually sets in, when, and whether it costs a studio anything measurable in retention or reviews. Twenty-two votes on two loops can't tell you. Neither can anyone else's clip test, including the ones that come out in your feed next month claiming otherwise.

The place to look next isn't another blind test. It's the middleware session. Ask a team using generated audio to open their Wwise or FMOD project and show you the state machine — how many tiers, how the transitions are quantized, what happens on the fail state. The answer to whether AI music generation belongs in games is sitting in that window, not in the render.

The question was never whether you can hear the difference in nineteen seconds. It's whether the music can hear you.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
I

Imogen Hale

Music-Tech & Licensing Reporter

Imogen Hale reports on the business side of AI music — licensing terms, royalties, and copyright — reading the fine print so working creators don't get burned. More by Imogen Hale →