Somewhere around the fourth pass, I noticed the hi-hat wasn't moving.
I was in a coffee place in the middle of a Tuesday afternoon, and the track on the ceiling speakers was a soft, warm, mid-tempo thing: a Rhodes chord with too much chorus on it, some vinyl crackle laid over the top like a tablecloth, a kick sitting politely under everything at around 82 BPM. It was fine. It was better than fine for what it was doing. And every eight bars the same hi-hat pattern came back with the same velocity on the same sixteenth note, the way a human drummer's arm never quite does, not even a bored one on the ninth take.
Nobody looked up. The person across from me was on a call. The barista was restocking oat milk.
That room is the whole argument about AI music ethics in miniature: something obviously generated, doing a job it did adequately, in front of people for whom the question never arose. And it's the reason the most popular rule about all this keeps collapsing when you lean on it.
The rule, stated as fairly as I can
Here it is, in the form you've probably heard: let AI have the background music, and leave the real music to humans.
It gets said from both directions, which is part of why it sounds so reasonable. Musicians offer it as a concession — fine, take the hold music, leave us the albums. People building the tools offer it as reassurance — nobody's replacing your favourite band, this is for the elevator. It has the shape of a truce, and truces feel like wisdom.
It's about two-thirds right. It's right about what these systems currently make well and wrong about how the people who make music actually get paid, which means it's comforting in exactly the place where comfort is least useful.
Let me take the two-thirds first, because it deserves more credit than the people repeating it usually give it.
Where the rule holds up
The bottom of the music market was industrialised decades before anyone typed a prompt into anything.
If you have never worked in it, the thing to understand is that a large amount of the music you hear in a day was never anybody's expression. It was inventory. Production libraries have been selling cues to advertisers, corporate video teams, radio stations and TV producers since long before streaming — catalogues with hundreds of thousands of tracks, tagged by mood and tempo, licensed on a buyout basis so the buyer pays once and never thinks about it again. The brief for that music was never "say something true." It was closer to: fill 47 seconds, sound optimistic but not distracting, no vocals, no drums in the first eight, give us a 30 and a 15 and a stinger, and don't get our client a copyright claim.
I have written to that brief. Plenty of composers I know still do. It is real craft — hitting a mood in eight bars with no room to develop is harder than it sounds — but it was already being priced, credited and consumed as a commodity. The name of the person who wrote the music under your bank's explainer video is on a spreadsheet somewhere and nowhere else.
So when someone says generative tools are coming for that tier, they are describing a market that had already been trained, for a generation, to treat music as a unit of stock. The coffee shop track wasn't displacing an artist's statement. It was displacing an anonymous cue from a catalogue with four hundred thousand siblings.
And the systems are genuinely competent at that specific job. Short loops, consistent mood, no vocals, endless variations until one fits the picture. If you need a bed that stays out of the way for a minute and doesn't modulate anywhere surprising, you can get one.
What they are still visibly bad at, as of writing, is anything that has to develop. Long-form structure falls apart: a four-minute generated track will often give you a chorus that arrives without the arrangement doing anything to earn it, a bridge that's the verse with a filter on it, an outro that stops rather than ends. Vocals remain the hardest thing in the room — sibilance that spits, consonants that smear, a breath placed where no lung would have put it. And the mixes tend to come back pre-glued: everything sitting in the same narrow stereo image, already squashed, with no headroom left for you to do anything about it. Ask for stems and you find out fast whether the tool you're using actually has them or is handing you a bounced master with a different name.
That's the honest version of the good news. On the job the rule describes, the rule is basically correct.
Where it starts to break: the boring work paid for the interesting work
The part the truce ignores is that the tier it hands over is the tier that funded everything else.
Almost nobody making music you'd call art makes their living from it directly. They make it from the layer underneath: session work, the wedding band, the two library cues a month, the jingle, the sync placement in a home renovation show, teaching on Thursdays. The record nobody paid for got paid for by the ad nobody remembers. This isn't a romantic claim, it's an accounting one, and it's how most working musicians I know have structured their lives for as long as I've known them.
There's a second thing the boring work does, which is teach. You learn arrangement by having to write eight versions of the same cue for a client who can't say what they want. You learn to finish things by having a deadline attached to somebody else's money. Remove the bottom rung and you don't get a shorter ladder, you get a ladder that fewer people can climb onto at all, and you find out about it in ten years rather than next quarter.
Here's the honest counterweight, because this argument gets overstated too: that ladder was already rickety. Streaming economics, the consolidation of sync buyers, the collapse of regional session work, the long slide in what a library cue pays — none of that was caused by generative models. If you want to be accurate, the tools arrived at a market that had been softening the ground for them for twenty years. Blaming them for the whole condition is a way of avoiding the older, duller story about how music got repriced.
But "it was already happening" is not the same as "it doesn't matter." It matters more, because the floor was already low.
What "you can hear the difference" actually means
The other load-bearing beam under the truce is the belief that people can tell. Push on it and it wobbles in an interesting direction.
You can often tell, and I can tell more often than most people, and the tells are almost never the ones civilians expect. They aren't about warmth or soul or some ineffable quality of the timbre. They're structural. A loop that repeats without a single variation across two minutes. A drum fill that lands perfectly on the grid and resolves into nothing. A chorus where the arrangement density doesn't change, so the section boundary is only implied and never felt. A vocal with no consonant grit. Lyrics that rhyme in the safest available direction every time, because the safest direction is what the model was optimising toward.
Now put that track at forty percent volume, on a phone speaker, in a kitchen, in the middle of a playlist you didn't build. Every tell I listed disappears. Structural failures need attention to register as failures, and functional listening is the opposite of attention. That's not an insult to the listener — it's what functional music is for.
And the uncomfortable part: humans make loops that never vary and drum fills that land on the grid, and some of that music is very good. I've watched people in a studio confidently identify a generated track that a person spent a week on, and confidently praise the humanity of something that came out of a prompt. Confidence runs about the same in both directions, which suggests the thing being detected is often the label rather than the audio.
So the rule survives here in a weakened form. Trained ears can usually tell in a listening room. The market almost never listens in a listening room.
The argument that isn't about the music at all
The loudest objection has nothing to do with whether the output is any good, and the debate keeps going in circles because two separate arguments have been wedged into one word.
The first argument is aesthetic: is this music worth anything. The second is about provenance: where did the capability come from, and was anyone asked.
Most of these systems learned what a chorus is by ingesting recordings. Which recordings, under what terms, and with what compensation to the people on them, varies enormously by provider and is frequently not disclosed in any checkable way. Some vendors advertise licensed or fully-owned training catalogues; those claims are worth reading closely rather than taking on trust, and the terms shift as litigation and licensing deals move. The legal picture differs by jurisdiction and is not settled.
What matters for the reader is that these two arguments are independent. You can think a generated cue sounds perfectly good and still object to how the model that made it was built. You can think the output is mush and have no objection at all to the training. Every time someone answers "they trained on my catalogue without asking" with "but the music sounds fine," they've changed the subject and usually not on purpose.
It's the difference between someone photographing your house, which is fine, and someone putting the photo of your house in an ad, which is a conversation you'd expect to be part of.
What each side is actually defending
It helps to stop treating this as one disagreement with two teams. There are at least five values in the room, each coherent, each wanting a different remedy.
| The value | What it's protecting | What would actually satisfy it |
|---|---|---|
| Craft | Skill built over years, and the standards that come with it | Nothing regulatory — this is settled by whether the work holds up next to other work |
| Consent | The right to decide whether your recordings train a system | Opt-in defaults, auditable training disclosures |
| Livelihood | The paid tier underneath the visible careers | Payment flowing to rights holders, or new floors under session and library work |
| Access | People with a deadline and no budget getting something usable | Cheap tools with licence terms that don't ambush them later |
| Provenance | Knowing what you're listening to and who made it | Disclosure at the point of listening, not buried in a terms page |
Read down that column on the right and you'll notice something: almost none of these remedies conflict with each other. A world with opt-in training data, clear labelling and cheap tools is available. Most of the heat in this argument comes from people arguing as though satisfying one value must cost another, when the actual fight is over who bothers to build any of it.
The access case, without the sermon
I want to state this one carefully because it usually gets delivered as a sermon and it doesn't need to be.
A solo game developer I know shipped a demo last year with a menu loop she made in about twenty minutes from a prompt. The relevant fact is not that she saved money on a composer. It's that there was never going to be a composer. The alternative was the temp track staying in the build, or silence, or the demo not shipping. That is what the bottom of the budget range actually looks like, and pretending otherwise is a way of arguing with a version of the situation that doesn't exist.
The counterweight is that the same collapsed price floor is what makes it less likely she'll hire a composer on the project after this one, when there is money. Both things are true at once. Neither cancels the other.
And there's a trap in the access case worth naming, since it's the one that burns people who assume free means free: the licence. Whether you can use generated output commercially, whether that survives you cancelling the subscription, whether it applies to a client's work or only yours, whether the provider indemnifies you if a rights claim lands — these differ between services and between tiers of the same service, and they change. The word "royalty-free" has never meant "no strings," not in stock libraries and not here. If you're going to use this stuff for anything that earns money, the licence page is the only page that matters.
How I'd decide what to be bothered by
If you're a listener rather than a maker, the useful move is to stop asking whether AI music is good or bad and start asking narrower questions that actually have answers.
Consent. Was the system built from work whose owners had a real choice? This is the question with the least available information and the most weight, and it's the one worth pushing publishers and platforms on.
Payment. Is money reaching the people whose recordings taught the model anything? Not in principle — in a scheme somebody can describe.
Disclosure. Do you know what you're hearing at the moment you hear it? Some platforms have begun labelling; coverage is patchy and inconsistent. You are allowed to want this without having any objection to the music itself.
Displacement. Was this replacing a person who was going to be hired, or filling a slot that had no budget? Different situations, different obligations, and conflating them is the most common error in both directions.
Job or claim. Was this music doing a job — filling time, setting a mood, keeping a room from feeling empty — or making a claim about a person's experience? The second is where the generated stuff still reads as counterfeit to me, and not for mystical reasons. It's that a claim about experience is checkable against a life, and there isn't one behind the render.
Who should worry about this, and who shouldn't
If your relationship to music runs through live shows, records with names on them, and a scene you can name three people in, you can put this down. That market has never been about audio efficiency and it isn't going to be. People go to shows to be in a room with someone. Nothing about generated audio touches that, and every prediction that it will has been made about every new instrument for a century.
If most of your listening is playlists built around a mood — focus, sleep, workout, dinner — then you're the person this actually affects, and the effect is that an increasing share of what you hear was made to fill a slot rather than to be heard. You may find you don't care. That's a legitimate answer and it's the one most people will land on.
If you're someone whose income came mostly from non-featured work, you already know. You didn't need this piece.
A more honest rule
The truce fails because it sorts music by category — background versus real — when the thing people are actually upset about sorts by provenance. Nobody is angry that a mood loop exists. They're angry about how the capability to make it was assembled, and about which tier of worker is absorbing the cost.
So the replacement rule I'd offer is duller and harder to say at a party: ask who consented and who got paid, not whether a human was in the room. A generated cue built from a licensed catalogue with money flowing back is a different object from an identical-sounding cue built from scraped recordings, even though your ears cannot tell them apart. That's not a failure of your ears. It's a sign that the question was never an audio question.
The myth is that AI takes the disposable music and leaves the real music to humans.
The truer version is that AI takes whatever music was already being bought by the yard, which happens to be the work that paid for the real music in the first place.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.