In April 2023, a track called "Heart on My Sleeve" put a synthetic Drake vocal over a synthetic Weeknd hook and pulled somewhere north of a few hundred thousand plays on Spotify before it was taken down. Within a week the framing had hardened: this is the end. Not the end of one loophole in voice likeness, not the end of a takedown policy — the end. Two and a half years later, that framing is still doing most of the work in conversations about AI in music production, and it's worth asking what it was ever built on.
Because it was built on something. Beliefs this durable usually have a source document. The interesting thing about this one is that the source is thinner than the belief it holds up.
The claim, stated honestly
The strong version goes: generative tools will produce music good enough and cheap enough that the people who currently get paid to make it will stop getting paid. Not "some workflows change." Displacement, at the level of a career.
The weak version — tools compress certain jobs, shift value toward the people who can direct them, and put real pressure on the lowest-margin corners of the business — is almost certainly true and much less interesting to write headlines about. Most of the industry conversation collapses the two. If you're a label deciding whether to greenlight AI-assisted production, or an artist deciding whether to touch it at all, the collapse is expensive: you end up making a strategic decision using a number that was never a forecast.
Where the number came from
The most-cited hard figure in this debate traces back to a study commissioned by CISAC — the international confederation of authors' and composers' societies — and carried out with Goldmedia, published at the end of 2024. Its headline finding, repeated everywhere since: music creators stand to lose roughly a quarter of their revenues to generative AI by 2028.
Read the study's own framing and two things become clear. First, it's a scenario model, not a measurement. It projects a market for AI-generated music, estimates substitution effects, and reports what falls out of the assumptions. That's a legitimate and useful thing to do — CISAC represents the people whose income is on the line, and modeling downside risk is exactly their job. Second, the substitution it models is heavily weighted toward the parts of the market where music is bought as a commodity: production music, library and sync placements, background beds. That's where the projected losses concentrate.
Neither of those caveats is hidden. They're in the study. They fall off in the retelling.
So the chain runs: a commissioned scenario model → a press release with a round number → a wave of coverage that drops the word "scenario" → a year of think-pieces citing the coverage rather than the model → a belief that feels like measured fact because you've encountered it forty times. Nobody lied. The provenance simply evaporated, the way it does.
The honest reading isn't "the number is wrong." It's that the number describes a modeled future for a specific slice of the business, and it's being used to describe the present for all of it.
This has a longer history than you think
In 1930, the American Federation of Musicians ran a national ad campaign against what it called "canned music" — the recorded soundtracks that were putting live pit orchestras out of movie theaters. The campaign had a mascot, a robot conductor called the Musical Robot. It lost. Something like tens of thousands of theater musicians lost steady work inside a few years. That displacement was real and it was permanent, and it's the strongest historical argument the pessimists have.
But look at what came next: the same recording technology that emptied the pits created the studio session economy, the film scoring stage, the record industry's entire session-player class. The jobs didn't come back. Different jobs appeared, in different places, for people with adjacent skills. Whether that's consolation depends entirely on whether you were 45 in 1932.
Run the same tape on the Linn LM-1 and the drum machine panic of the early eighties — the "drum machines have no soul" bumper stickers were a real thing session drummers printed. Session drumming contracted. It didn't vanish; it stratified. The top of the market got more valuable and more specialized, the middle got automated, and a new job (programming convincing drum parts) appeared that paid people who understood drums.
Auto-Tune ran the cycle again in the 2000s, and the pattern held: universal condemnation, then quiet universal adoption, then a generation of producers using it as a deliberate texture rather than a repair tool.
The pattern in all three: the predicted catastrophe (the craft dies) doesn't land. The unpredicted one (the middle of the market gets hollowed out, fast, while the top gets more concentrated) does. That's not reassuring. It's just differently shaped than the headline.
What the tools actually can't do, as of writing
I've spent enough sessions with the current generation of generators to be specific about where they fall down, and the gap between demo and deliverable is still wide.
Vocals remain the hardest thing. Sustained notes wobble, consonants smear, and anything with a plosive tends to arrive with a strange metallic edge on the transient. Long-form structure is worse — ask for a four-minute piece with a real development section and you'll typically get two good ideas and ninety seconds of filler between them.
Stems are the practical dealbreaker for professional work. Some tiers hand you a stereo bounce, which means you cannot duck the bass under dialogue, you cannot mute the lead for a version, and you cannot deliver the alt mixes a client will ask for. If a tool doesn't give you separated stems at 48kHz, it isn't in your pipeline for paid picture work regardless of how good it sounds.
And prompt-roulette is real. You can burn forty minutes on renders trying to get a specific thing — say, a detuned analog bassline sitting under a broken 808 at 92 BPM in F minor — and end up with twelve near-misses and one usable eight bars. The time saved is real but it's lumpy, and anyone selling you a smooth productivity curve hasn't shipped a cue on a deadline.
Where the pressure actually sits
Here's the part the existential framing gets backwards. The threat isn't distributed evenly across music, and it isn't aimed at the thing most people are defending.
An artist with an audience sells a relationship. Generative tools don't compete for that; nobody forms a parasocial attachment to a prompt. What generative tools compete for is music bought as a commodity input — the sixty seconds of tense-but-not-too-tense underscore for a corporate explainer, the royalty-free bed for a YouTube channel, the eight-bar sting for a podcast transition. That market has been price-compressing for fifteen years already, since long before any of this. Generative tools are accelerant on a fire that stock libraries lit.
If you write production music for a living, the pressure is here now and the CISAC model is describing something close to your reality. If you're an artist with listeners, the honest answer is that the near-term threat to your income is the same one it was in 2019: streaming economics, playlist gatekeeping, and touring costs. AI didn't change your math nearly as much as the discourse suggests.
The cruelty is that the production-music floor is where a lot of working musicians make rent while building the artist career. The ladder is being pulled up in the middle, not the top. That's a specific, addressable problem, and it gets almost no attention compared to the abstract one.
So: should we use this?
The question decomposes into four smaller ones that actually have answers, and none of them are about whether AI is good or bad.
What was it trained on, and will they tell you in writing? Some vendors license catalogs and say so. Some are vague. The vagueness is the signal.
Is there indemnification, and what does it cover? Read whether it covers you or covers the vendor, and whether it survives a change in your subscription tier. This varies enormously between tools and changes often — check the current terms rather than what someone told you last year.
Does any voice in the output belong to a person who agreed to it? Imogen Heap's approach — building a licensed model of her own voice that she controls and can revoke — is the version of this that respects consent. Cloning a living singer who didn't agree is a different act entirely, and no amount of "it's just a tool" collapses that distinction.
Will you say you used it? Not because disclosure is legally required in most contexts — it generally isn't — but because the reputational damage in every case so far has come from the denial, not the use. Artists who said up front what they did have absorbed it. Artists who denied and were caught have not.
That's what "responsible adoption" reduces to. Not a position on the technology. Four questions with paper trails.
Try this before you decide anything
Take one real brief you've already been paid for — a cue you delivered, with the client's actual notes — and run it through a generator you're evaluating. Not a fresh creative idea; a job you know the shape of. Time it honestly: prompt iterations, render waits, the editing you'd still have to do to make it deliverable, the stem situation.
Then compare that number to what you actually billed. You'll get one of three answers: the tool saves you real hours on this kind of work, it doesn't, or it does but only for the parts of the brief you never enjoyed anyway. All three are useful, and all three are yours rather than someone's model.
The existential threat is a story about the future. The invoice is a fact about your week.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.