Home/ The Signal/ Industry/ Can AI Music Detection Tools Keep Your Playlists Human? I Tested Them On My Own Library
Playlists

Can AI Music Detection Tools Keep Your Playlists Human? I Tested Them On My Own Library

The track was 74 BPM, minor key, the kind of warm neo-soul instrumental that ends up three songs deep in a focus playlist. I liked it for about forty seconds.

A photorealistic portrait of a young woman in her late twenties seated in a…

The track was 74 BPM, minor key, the kind of warm neo-soul instrumental that ends up three songs deep in a focus playlist. I liked it for about forty seconds. Then I noticed the hi-hats: every single one landing at identical velocity, no drift, no ghost notes, nobody's wrist getting tired. The room tone underneath was a room that has never existed. So I did the thing the internet now tells you to do and ran it through one of the free AI music detection tools that have appeared over the past couple of years.

Here is the verdict before anything else: these scanners are genuinely good at catching mass-produced, fully generated filler, and genuinely bad at telling you whether a person sat in a room and made a decision. They answer a narrower question than the one you are actually asking.

The question underneath the question

When you ask "is my playlist AI," you are almost never asking a technical question. You are asking whether the money you pay every month is reaching anyone. You are asking whether the thing that moved you on the drive home was made by a person who also once got moved by something on a drive home.

A classifier cannot answer that. It can answer a much smaller question — does this waveform statistically resemble the output of a known generator — and that smaller question turns out to be useful anyway, as long as you know which one you are getting back.

What these scanners are actually listening to

Not lyrics. Not vibes. Not whether the chord progression is boring, which would flag most of commercial radio.

Generative audio models leave consistent statistical residue in the signal: characteristic behavior in the high frequencies, in stereo width, in how transients are reconstructed after the model's own compression stage. Detectors are trained as classifiers on large piles of known-generator output set against human recordings, and they learn those fingerprints rather than anything musical.

Which tells you exactly where they break. A model trained on the output of the generators that existed when it was built is sharpest on those generators. Feed it something from a newer system, or something that has been bounced to a 128kbps stream, run through a mastering chain, and re-encoded twice on its way to your ears, and the residue it is hunting for gets smeared. Vendors publish accuracy figures on their own held-out test sets. Those figures are real and also not the conditions you listen in.

How I'd decide which scanner to trust

What it can see. Some tools take one uploaded file. Others connect to your streaming account through OAuth and walk your saved library. The second is far more useful and asks for far more — you are handing a third party read access to your listening history. Check that you can revoke it from your streaming service's app settings, and actually do it when you're done.

What it claims to detect. The honest ones say "fully AI-generated" and mean it. A track where a human wrote the parts and used a model for mastering or stem separation is not what these are built to catch, and a tool that implies otherwise is overselling.

What it does with the audio. Does your upload get retained as training data? This is buried in terms of service and it is the single thing most worth reading.

How it reports uncertainty. A scanner that returns a confidence band and admits to a maybe is more trustworthy than one that returns a binary verdict on everything. Certainty is the tell of a tool that hasn't been tested hard.

Price over twelve months. Most of the consumer-facing ones are free right now, offered by streaming services and detection startups building credibility. Free things that require an account tend to acquire a price. Assume the free tier narrows.

What happened when I ran my own catalog

I scanned a saved playlist of roughly two hundred tracks, mostly ambient and instrumental hip-hop, which is the genre where this problem is worst because it is the genre where uploads are cheapest to generate.

A close-up photograph of a professional recording studio drum kit at rest, focused tightly…

The results were reassuring in the boring way. The overwhelming majority came back clean. The flags clustered exactly where you'd expect: single-track "artists" with no bio, no live dates, five hundred monthly listeners and thirty albums, all of them beat tapes with names like Rainy Study Session Vol. 14. Nothing I had a relationship with got flagged. Nothing anyone would call a favorite record.

Then I ran my own work through it, which is where it got interesting. A batch of loops I cut for a game score in 2023 — hardware-sequenced, played by my hands, quantized hard because the engine needed sample-accurate loop points — came back as possible. Not confirmed, but not clean either. Grid-perfect timing and a clean 48kHz bounce apparently looks a lot like a machine, because in the ways the classifier measures, it is a machine. My old indie record from a decade ago, tracked in a bad room with a lot of bleed, sailed through without a flicker.

That is the useful lesson. These tools reward mess. If you make electronic music, library music, or anything mixed to a modern loudness target, expect the occasional false positive on genuinely human work, and do not treat a flag as an accusation.

Where the answer is honestly "it depends"

There is no clean line, because production hasn't had one for years. Consider the spectrum a single track can sit on: fully prompt-generated with no human input; human-written with an AI-generated vocal; human-performed with an AI mastering chain; human-recorded and rescued with AI stem separation; sample-library instruments a person arranged with care.

Detectors are built to catch the first case. Case two is contested. Cases three through five are how essentially all commercial music is made now, including by people you would defend at a dinner party. If your standard is "no model touched this at any point," no scanner will enforce it, and honestly, neither will the credits.

The platforms do not agree with each other

This is the part that matters if you are switching services. Deezer has been the most public about the scale of the problem, reporting that a large and rising share of tracks delivered to it daily are fully generated, and it labels albums it identifies as such. Spotify has moved toward industry-standard AI disclosure credits carried in track metadata, which depends on labels and distributors filling that metadata in honestly. YouTube leans on uploader disclosure. Smaller platforms lean on human curation and a smaller catalog.

So the same track can be labelled on one service, silently present on another, and absent from a third. Switching platforms means switching definitions of what counts as disclosed, not switching into a clean catalog.

Who should run a scan, and who should skip it

Run one if you're building something that carries your name — a wedding playlist, a podcast bed, a set you're getting paid for. Run one if a specific artist you support feels off and you want to check. Run one out of curiosity, once, on your most-played playlist, because knowing the actual proportion is better than imagining it.

Skip it if you can feel yourself about to audit every ambient record you own. The scan will hand you ambiguous results on borderline material and you will spend an evening prosecuting a musician who spent three weeks on a track. The proportion of generated material in a normal personal library is currently low. The proportion in algorithmic mood playlists is where the pressure actually is, and that is a curation problem you fix by following artists instead of moods.

What we have right now is not really a detection problem. It is a labeling problem with a detection workaround bolted on while the labeling catches up. Scanners are the interim measure — imperfect, free for now, and worth the ten minutes.

Use the scan to clear the filler. Then trust your ears, because they are the only detector nobody can train around.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
I

Imogen Hale

Music-Tech & Licensing Reporter

Imogen Hale reports on the business side of AI music — licensing terms, royalties, and copyright — reading the fine print so working creators don't get burned. More by Imogen Hale →