Home/ The Signal/ Industry/ AI Music Generation Detection, Tested: What Treblo and Sonauto Can Actually Prove
Authenticity

AI Music Generation Detection, Tested: What Treblo and Sonauto Can Actually Prove

The thing that bothered me was not the voice. It was the stereo width. Somewhere in the second verse of the track half my feed argued about this summer, the arrangement stacks a counter-melody and a…

A photorealistic environmental portrait of a sound engineer seated in a home studio at…

The thing that bothered me was not the voice. It was the stereo width.

Somewhere in the second verse of the track half my feed argued about this summer, the arrangement stacks a counter-melody and a layer of claps, and the image does not move. Same width, same depth, same air on top as the first verse. Real mixes lurch a little when you pile things on; something masks something else and an engineer chases it around with an EQ. This one sat perfectly still, like a photograph of a mix. That itch is what sent me into AI music generation detection — first Treblo, the app whose detector set the argument on fire, then Sonauto, the model shop behind it, to find out whether I could get the machine to hand me the same song back.

The verdict, up front, because it is the part everyone skips: a detector can tell you a track is consistent with the output of a specific model, which is genuinely useful, but it cannot tell you that no human touched the file. The only evidence that has ever moved me from suspicion to something I would say out loud is reproducibility.

The question you have probably already asked

You have had this thought with headphones on. Can you actually prove a song was made by AI?

The honest answer starts with a definitional problem. Almost every record you like was made by a chain of processes, some of them automated for twenty years. Melodyne rewrote pitch. Drum replacement swapped the snare. A stock loop pack supplied the shaker. Nobody calls that synthetic. What people mean when they ask the question is narrower and more emotional: did a person write this, or did a person type a sentence and pick a take?

Detection tools do not answer that question. They answer a smaller one — does this waveform look like it came out of a particular generator — and then the internet does the emotional math on top.

What Treblo and Sonauto actually are

Treblo is the consumer-facing side: type a description, get a full song with vocals in a couple of minutes. Sonauto is the model shop underneath it. That relationship is the single most important fact about the detector, and it cuts both ways.

Detecting your own model's output is a fundamentally easier problem than detecting "AI music" in the abstract. If you built the generator, you know its artifacts from the inside: the way it resolves consonants, the spectral signature it leaves above 16 kHz, the fingerprints in how it renders reverb tails. It is the difference between recognizing your own handwriting and identifying an arbitrary forger.

So a positive result from a first-party detector is a meaningful statement about that family of models. A negative result is close to meaningless about everything else. If a track came out of some other generator, or out of an open-weights model someone fine-tuned in a bedroom, you should expect a shrug.

How I tested it

This is a producer's bench test, not a study. I have no lab, no labeled corpus of thousands, and no access to anyone's weights. Here is what I actually judged, so you can weigh it accordingly.

  • Reproducibility. Can I sit down with the generator and produce something that lands in the same neighborhood — tempo, key, arrangement shape, vocal character — using an obvious prompt?
  • Scope. Does the tool say which models it claims to recognize, or does it imply universal coverage it cannot have?
  • False positives. I fed it material I have the session files for, including two cues I wrote for a short film and one embarrassing 2019 synth-pop demo of my own.
  • Legibility. Does it hand back a probability with context, or a verdict-shaped badge that will get screenshotted into an argument?
  • Independence. Who benefits from the answer, and has anyone outside the building checked the method?

What the machine predictably gives you

The reproducibility test is the one worth your evening. I typed something deliberately generic:

synth-pop, 112 BPM, A minor, male vocal, sidechained analog bass,
tape-saturated chorus lift, no bridge, radio edit

What came back, across several renders, shared a grammar. Eight-bar intro. Vocal enters on the nine. First chorus lands somewhere around forty-five seconds, which is exactly where a playlist skip-test wants it. A doubled chorus vocal, panned wide, mixed a hair forward of where a human engineer would leave it. No bridge, or a "bridge" that is the chorus with a low-pass filter on it. And endings that fade, because a model that has learned song shape from streaming audio has learned that songs mostly stop being played rather than stop.

The sonic tells were subtler and more interesting. Sibilance smears — the s on a held note spreads instead of snapping. Cymbals arrive as a wash rather than as struck metal with a decay you can follow. The bass sits so precisely on the grid that there is no pocket at all, none of the microscopic pushing and dragging that makes a groove feel like two people agreeing.

A photorealistic studio photograph of a professional audio mixing console in a dim recording…

Here is the counterweight, and it matters: I have heard human bedroom-pop that does every one of those things. Producers learn from the same reference tracks, quantize to the same grid, and use the same three chorus-lift plugins. "It sounds like a machine made it" is a hypothesis, not a finding. That is precisely why the reproducibility step exists.

Where the answer turns into "it depends"

This is the section that keeps me from writing a cleaner article.

Stems and re-sings. Export the generated instrumental, put a real vocalist on top, print it through outboard gear, and you have a hybrid record. What should a detector say about that? What should a label say about it? There is no settled answer, and the tools are not built to express one.

Scratch demos. Plenty of working writers now generate a rough sketch to hear an arrangement idea, then rebuild it by hand. If a few seconds of the original bounce survive into the master, a detector may light up on a song that a human genuinely wrote.

Mastering as laundry. Heavy limiting, saturation, and a lossy encode chew on exactly the high-frequency detail that classifiers lean on. A track that flags clearly at 48 kHz WAV can get quieter, evidentially speaking, after it has been through a streaming codec twice.

Model drift. Generators update. Detectors chase. A tool that is accurate against last quarter's checkpoint is making a weaker claim about this quarter's, and neither the vendor nor you can always tell which one produced a given file.

Silence is not proof. An artist who declines to post session files is not thereby guilty. Session files also prove less than people think: a project can be reverse-built around a finished bounce by anyone patient enough.

One honest negative each

Treblo. The detector is run by a company with a stake in the answer, and as of writing I have not seen a published methodology, a stated false-positive rate, or an outside review of the classifier. That does not make the results wrong. It does mean a screenshot of a confidence score is an assertion by an interested party, and it deserves the same scrutiny you would give any other press release. The results I got were also binary-feeling in presentation — the sort of output that travels through a quote-tweet with all the nuance stripped off.

Sonauto. Vocals remain the weak joint. Dense arrangements — layered guitars over a busy kit — went mushy in the midrange in a way no amount of prompt fiddling fixed, and consonants degraded first. On licensing, read the terms on the tier you are actually paying for rather than the marketing page, and read them again when you renew. Commercial-use language on generative audio tools has changed more than once across the category, and it is the single place creators get burned worst.

Labels, metadata, and the part platforms have not solved

Several distribution platforms have started asking uploaders to disclose generative content, and legislators in more than one jurisdiction have floated labeling requirements. Treat all of it as a claim rather than a fact. Disclosure is self-reported. Embedded provenance metadata is fragile — bounce a file through a DAW, re-upload it from a phone, and the tag is gone. Content-credential standards that took root on the image side are creeping toward audio, but a credential proves what a cooperative tool wrote, not what an uncooperative one omitted.

Who this is for, and who should skip it

Use a detector if you are about to attach money or a byline to a claim: a sync supervisor clearing a track, an A&R person about to sign someone, a journalist who needs more than an ear. Use it as one input, and pair it with the reproducibility test.

Skip it if you want ammunition for a comment thread. Skip it if you are hoping for a universal detector, because nobody has one. And skip it especially if you are an artist trying to prove your own innocence — these tools are built to flag, not to exonerate, and a clean result will not convince anybody who already decided.

The part nobody really wants

The strangest thing about running these tests was how little the answer changed anything. I played my reproductions for a friend who had the original on repeat all July. He listened, said "huh," and kept the song in his playlist. He was not being dishonest. He had made a distinction I keep failing to make cleanly: he cared whether the song worked on him, not who was in the room when it was made.

That is a legitimate position, and it is worth separating from the other one going around, which is "if I do not check, it is not true." The first is taste. The second is a decision to stay comfortable, and it is the reason detection tools will keep mattering more to the industry than to the audience.

So, the rule you can use tonight: if you cannot sit down with the tool and get it to hand you the song, you have a suspicion, not a finding.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
J

Juno Park

Game Audio Writer

Juno Park covers AI sound design and game audio workflows — foley, loops, and middleware — after seven years cutting assets for mobile and indie titles. More by Juno Park →