A perfect AI voice cloning detector would not have stopped the call that emptied your mother's account.
I want that stated flat at the top, because the rest of this piece is spent earning the right to have said it. Detection is where most of the money and nearly all of the policy attention have gone — classifiers that take a clip of speech and hand back a probability that a machine produced it. Those classifiers keep improving. The fraud keeps getting worse anyway, and not because the engineering behind the detectors is lazy.
I repair and mix audio for a living. Most weeks that means chasing a hum out of a dialogue take, or deciding whether a synth pad survived a bounce to a lossy format. It has made me irritatingly literal about one thing: audio is not a message, it is a signal, and a signal changes shape every time it moves through something. That, more than any argument about model capability, is why detection loses.
Three seconds, and what three seconds buys
Vendors advertise usable clones from a few seconds of clean speech. Three seconds is the number that keeps surfacing in demos and product copy, and for a quiet, close-mic'd sample it is at least directionally honest. Three seconds is a voicemail greeting. It is a toast at a birthday someone posted. It is a church announcement, a school board comment picked up by the local cable feed, four words from a podcast guest spot.
The asymmetry there is the whole problem. Building the weapon costs one clean breath of audio and the price of a monthly subscription. Defending against it, under the current design, costs a real-time analysis pipeline attached to every phone call in the country. Nobody has built that, and if they did, it would not work — for reasons that have nothing to do with the classifier and everything to do with what a phone line does to sound.
What a detector is actually listening for
A detector does not hear "fake." It looks at fine structure that synthesis tends to get subtly wrong: over-smoothed spectral envelopes, a noise floor that stays unnaturally consistent between phonemes, breath that arrives on a schedule, high-frequency energy between harmonics that behaves too politely. A real larynx is a chaotic mechanical object. Vocoders are not, quite, and the residue shows up in the top of the spectrum and in the microsecond timing of glottal pulses.
Classifiers trained on spectrograms find that residue reliably — on the generator they were trained against, at full bandwidth, in a file. Three conditions in that sentence, each one load-bearing.
Can you detect an AI voice on a live phone call?
Not reliably, and not in time. A phone call is not a recording. It is a heavily compressed, band-limited reconstruction of one. Traditional telephony carries roughly 300 Hz to 3.4 kHz, sampled at 8 kHz; mobile and VoIP legs re-encode again through their own codecs. Everything above that ceiling — where a large share of synthesis artifacts live — is discarded before any analyzer gets a vote. Then automatic gain control flattens dynamics, noise suppression scrubs the room tone, and packet-loss concealment literally invents short stretches of waveform to paper over dropouts.
Put a detector on that and you are asking it to identify a forgery from a photocopy of a photocopy. Worse, the same processing chain mangles genuine speech, which pushes false positives up exactly when you tighten the threshold.
| Where you'd run detection | What it gets to hear | Why it slips |
|---|---|---|
| Live, carrier-side, during the call | Narrowband, codec-quantized speech | The artifacts it needs were stripped in transit |
| A saved voicemail | Same, plus another round of re-encoding | The verdict lands after the callback already happened |
| Uploaded file, forensic analysis | Full-band original, if one exists | It usually doesn't, and the money has moved |
| Provenance marker added at generation | Whatever the generator embedded | Covers cooperating tools only; survives some processing, not all |
The verdict arrives after the decision
Suppose the score existed and were accurate. Deliver it to whom, and when?
These calls are engineered as timed social exploits. A crash, an arrest, a lawyer who needs a retainer in the next hour, a request not to tell the rest of the family. The window between first contact and irreversible transfer is often shorter than the time it takes a bank branch to open. A confidence figure of 0.71 arriving in that window, to a person who believes they are hearing their child, changes nothing. People who can recite every rule about not trusting unsolicited callers still comply, because the attack does not target their knowledge. It targets the part of them that responds to a familiar voice in distress.
And the failure mode runs both directions. Flag enough real, frightened calls as synthetic and you train families to dismiss the flag — which is a worse outcome than never having shipped it.
Even the good-faith safeguards are attributive, not preventive
The serious cloning vendors have spent genuine money here, and it should be said plainly: consent capture at enrollment, moderation queues, internal classifiers that recognize their own model's output, provenance markers embedded at generation. That is not a rounding error of effort, and some of it is technically hard.
It is also, almost entirely, backward-looking. It answers "did this audio come from us" after someone hands over the file — useful for an investigation, irrelevant to the sixty seconds in which a person decides whether to drive to the bank. The consent step has a sharper problem: on many platforms it means recording yourself reading a supplied sentence, which is a task a clone performs well. Consumer Reports has evaluated consumer cloning products and found the safeguards on several amounted to an attestation checkbox. And every one of these controls applies only to hosted services. Open-weight synthesis models run on a laptop answer to no terms of service at all.
Why the market ships detectors instead of brakes
This is where I stop blaming engineers. A detector is a product: it has buyers, a price, a dashboard, an audit log a compliance team can point at. Identity verification at enrollment is not a product — it is friction, measured in lost signups, applied to a company's own funnel. Voluntary provenance marking has the structure of a commons: it works if everyone adopts it and pays off best for whoever quietly doesn't.
So the safeguards that get built are the ones that land on the revenue side of the ledger, and the ones that would actually reduce the supply of cheap cloned voices land on the cost side. No villain is required for that outcome. It follows from the incentives, and it will keep following from them until something external changes the arithmetic.
Most of the generative audio tools we cover at City of Punk are aimed at scoring a level or a trailer, and the consent-and-provenance language usually sits in the same section of the terms as the commercial-use grant. Read both halves; they tell you what a company thinks it owes.
The verification protocol that actually works
None of this requires technology. It requires a rule agreed on in advance, while nobody is frightened.
- Hang up and call back on a number you already have saved. Not a number the caller gives you. You should reach the actual person, or their real voicemail, within a minute — and if they answer confused, you have your answer.
- Agree on a passphrase now. Something absurd and unpublished. It works when the caller produces it without prompting; a caller who deflects has told you everything.
- Ask a question the internet cannot answer. Not the trip that got posted. Something like what the dog was called before you renamed it.
- Treat the payment method as the tell. Wire transfers, gift cards, crypto, or a courier arriving for cash are the signature of this crime, whatever the story attached to them.
- Put friction on the account in advance. Ask the bank in writing for a callback requirement on transfers above a threshold and a trusted contact on file. You'll know it took when they confirm it in writing.
- Reduce the training data where it's cheap. A default voicemail greeting instead of a recorded one, tighter privacy on public video. This is surface reduction, not protection.
What a rule would have to ask for to matter
Mandating detection regulates the losing half of the exchange. The measures with a plausible mechanism sit earlier or later in the chain: verified, liveness-checked consent before a voice can be enrolled; durable provenance obligations at the point of generation rather than after distribution; and, at the far end, payment-rail delay windows and liability placed on the party actually positioned to build friction. Banks and card networks solved a comparable arms race not by identifying every fraudulent transaction in real time but by making reversal and delay normal.
I'm describing options, not predicting outcomes, and none of this is legal advice. Any such rule reaches hosted providers and leaves locally-run models untouched. That is not an argument against it. It is an argument for being honest about what it buys: a raised floor, not a wall.
Three seconds is what the attack costs to build. The defense that works costs about thirty — hanging up, dialing the number you already had, and letting a real phone ring.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.