Three seconds. That is the number on the box — the length of reference audio a current voice-cloning model claims it needs to speak in a specific person's voice. It is the number doing the quiet work behind every briefing you have read about AI vishing attacks on asset managers, and it is the number your risk committee will quote back at you. It deserves a harder look than it usually gets, because of the three questions you actually need answered, it answers one.
I clone voices for a living, in the boring way. I score indie games and short films, and when an actor is unavailable for four lines of pickup dialogue and the shoot is closed, I build a model from their session audio and generate the pickups. So I have spent a lot of hours listening to what these systems get right and — more useful to you — precisely where they come apart. What follows is not a threat briefing. It is a description of an instrument, written by someone who plays it.
What the number measured
Three seconds is enrollment audio: the sample the system uses to build a speaker embedding, a compact numerical description of a voice. That embedding captures the static properties of a vocal tract — formant structure, fundamental pitch range, a couple of articulation habits like a soft R or the little intake of breath someone takes before a consonant cluster. Feed the embedding to a synthesis model with a script and you get speech that sounds, to a listener not braced for it, like that person.
That is what the figure was measured against: a short scripted read, judged for "is this the same person," produced from clean source. One speaker. Close mic. No crosstalk. Wide bandwidth. Every one of those conditions is load-bearing, and the honest framing is that this is a vendor-shaped number — a capability claim from the people selling the capability, not a finding about what happens on a Tuesday afternoon call to your operations desk.
Usually, when you strip a lab condition away, the result gets worse. Here it gets worse for you.
The phone line is working for the attacker
A voice call is not hi-fi and has never pretended to be. Traditional telephony carries roughly 300 Hz to 3.4 kHz; mobile codecs do their own aggressive throwing-away of anything the intelligibility model considers surplus. What is surplus, as far as a codec is concerned, is exactly where synthetic speech leaves fingerprints: sibilance up above 8 kHz, breath noise between phrases, plosive energy, the consistent low-level room tone of an actual human sitting in an actual office.
In a 48 kHz WAV I can hear a clone inside a bar. Downsample that same file to narrowband telephone quality and most of my tells are gone — not concealed, deleted. The channel your firm uses for its most sensitive verbal instructions has been discarding evidence for a century, and it does not know whose side it is on now. Anything sold to you as "listen for the fake" is being asked to work in the one medium engineered to destroy what it needs.
What three seconds does not buy
Here is where the gap actually sits, in my experience building these things:
| The clone handles this well | The clone still fumbles this |
|---|---|
| Timbre — sounds like them within one sentence | Prosody on an unscripted answer |
| Pitch range, accent, cadence on a read script | Turn-taking: overlap, barge-in, "no, sorry, go ahead" |
| Filler words, if trained on enough material | Latency — a beat too long before an off-script reply |
| A confident, urgent, well-rehearsed delivery | Shared history: last week's dinner, a running joke |
| A single sustained emotional register | Room consistency across a six-minute call |
Timbre is static and therefore cheap to model. Conversation is dynamic, two-way, and on a clock. The gap an attacker has to close is not identity — that part is solved and inexpensive. It is improvisation under real-time pressure. Which is precisely why the calls that work are tightly scripted and manufacture urgency: urgency is what keeps the target from generating the unscripted turn that breaks the model.
The reporting that put this on your agenda named household-name multi-strategy funds, and the peer-level specificity is the point. Nothing about the enrollment cost curve makes you a harder target than they were.
How I'd size your exposure
Five criteria, in the order I would work them:
Enrollment supply. Count the hours of each senior voice that exist publicly — earnings calls, conference panels, podcast appearances, the CFO's forty-minute fireside chat on a conference YouTube channel. The three-second figure means the marginal cost of adding one more target is effectively zero, so the question is not whether your executives are clonable. It is how far down the org chart public audio goes.
Decisions that accept a voice as an authorizing factor. Write the actual list: wire release, vendor bank-detail change, MFA reset, prime broker instruction, privileged access grant, fund admin exception handling. If a voice alone can move any of those, that item is your surface.
Callback discipline. When your staff calls back to verify, whose number do they use — the one in your directory, or the one that just called them, or the one in the email signature? This single question separates firms that blocked these calls from firms that did not.
Time-of-day coverage. These calls land at 5:50pm on a Friday, at the junior on the rotating on-call, during a holiday week. Test your control where it is thinnest, not where it is staffed.
Whether refusing is safe. If your culture penalizes the analyst who made the COO wait eleven minutes for a callback, your protocol is decorative. The control that matters is whether a junior person can impose friction on a senior voice without professional cost.
The defense is a protocol, not an ear
Detection tooling is real and improving, and I would still not build a program on it. It has to run on narrowband audio, which is the hardest possible input, and it fails in the expensive direction: a detector that flags two percent of legitimate executive calls gets switched off within a month, and everyone remembers it as the thing that cried wolf.
What survives contact is dull and structural. No voice-only authorization for money movement or credential changes, without exception and without a seniority override. Out-of-band callback to a directory number your firm controls. A short rotating shared phrase for high-value instruction paths. A mandatory delay on anything urgent — attackers cannot tolerate a two-hour clock, and legitimate business almost always can.
Train staff on the protocol rather than on the audio. Ear-training against synthetic voices produces a skill that decays, that the codec undermines, and that gets worse every quarter as the models improve. A callback rule does not degrade.
Who should act tonight, and who can stand down
Act tonight if a voice can trigger a wire or a credential reset anywhere in your stack, if your ops coverage rotates through junior staff outside market hours, or if your executives publish long-form audio — which is most of you, since that audio is a marketing asset nobody is going to retire.
You can reasonably deprioritize this if you already run dual control with out-of-band callback on one hundred percent of instruction changes. Your existing control absorbs cloned-voice fraud as a side effect. Spend the money on your exception paths instead, because the exceptions are where every one of these calls is aimed.
One honest limit: I am a sound person, not your CISO. I can tell you what the instrument does and where its seams are. Turning that into a controls framework is your job, and anyone selling you a single product as the answer is selling you a brochure.
If a voice asks you to move money, change a bank detail, or reset access, hang up and call back on a number you looked up yourself — and treat any reason you are given not to as the attack itself.
Not sure which tool to use?
Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.