Home/ The Signal/ Tutorials/ How to Build a Help Desk Script That Survives AI Voice Cloning Attacks
Security

How to Build a Help Desk Script That Survives AI Voice Cloning Attacks

Stop training your help desk to hear the fake. They cannot hear it, and neither can I, and I have spent a decade listening to audio for a living.

A tight, moody photorealistic close-up of a vacant help desk workstation at dusk: a…

Stop training your help desk to hear the fake. They cannot hear it, and neither can I, and I have spent a decade listening to audio for a living. The defense against AI voice cloning attacks is not a better ear and not a detector bolted onto your call-recording stack. It is a script — seven steps a tier-1 analyst can run at 4:50 on a Friday while a very convincing CFO is shouting about a wire that has to move tonight. The rest of this piece earns that claim and then hands you the script.

The news peg, briefly: over the past couple of years, business press has reported coordinated rounds of vishing calls hitting several large hedge funds and asset managers inside a short window — the same pretext, different firms, same week. Coordination is the interesting part, not the technology. The technology is old news. A German energy firm's UK subsidiary lost a wire in 2019 to a synthesized executive voice, reported at the time as a novelty. It stopped being a novelty when cloning became a consumer feature and the marginal cost of the fortieth call dropped to roughly the cost of the first.

Why the phone line deletes the evidence

Here is the part I can speak to from the studio rather than the SOC. A traditional phone call is not audio in any sense a mastering engineer would recognize. Narrowband telephony carries roughly 300 Hz to 3,400 Hz, sampled at 8 kHz, encoded with G.711 or a mobile codec at a fraction of the bitrate of a podcast file. So-called HD voice widens that to about 16 kHz sampling and it is still nowhere near a 48 kHz WAV.

Now consider where synthesized speech actually falls apart. Sibilance artifacts — the smeared, slightly metallic quality on an "s" or "sh" — live between 6 and 12 kHz. Breath noise before a phrase is broadband and quiet. Room tone, the low rumble of an actual office, sits under 100 Hz. The telephone band strips every one of those before the sound reaches the analyst's headset. The forensic evidence is filtered out by the transport, not hidden by the attacker.

It gets worse. When a codec detects silence it stops transmitting and the receiving end generates comfort noise — the network fabricates a plausible hiss so the line does not sound dead. The "background office" your analyst finds reassuring may be a synthesizer at their end of the call. And the one cue that does survive, prosody, is the cue attackers deliberately mask: cloned speech gets rhythm and emphasis subtly wrong, which is indistinguishable from a stressed executive on a bad connection in an airport. The panic in the pretext is not only social engineering. It is noise reduction.

Can you tell if a voice on a phone call is AI-generated?

No, not reliably, and you should design as if the answer is permanently no. Trained listeners perform poorly on narrowband, codec-compressed speech, and every improvement in synthesis narrows the gap further while the phone network stays exactly as lossy as it was in 1990. Treat voice as an identifier — a claim about who is calling — and never as an authenticator. Everything below follows from that one reclassification.

The target is your tier-1 analyst, not your CFO

The executive is the costume. The vulnerable process is the identity help desk, because that is the one desk in the building whose entire job is to restore access to people who have lost it. Password reset, MFA device re-enrollment, phone number change on the account of record: three actions that convert a convincing voice into a valid session. Your attacker does not need the CFO's money. They need the CFO's authenticator moved to a phone they own.

That desk is also measured on handle time, staffed junior, and trained on empathy. It is the softest, best-instrumented target in a financial institution, and it is defended by procedure or not at all.

The seven-step script

Write this into the ticket macro, not the wiki. If it is not in the workflow the analyst is already clicking, it does not exist.

1. Name the three actions that are callback-only. Password or passphrase reset, MFA device enrollment or reset, and any change to a payment or contact detail of record. Worked when: selecting any of those three in the ticket tool refuses to advance and opens the verification path instead.

2. Pull the callback number from the directory, never from the caller. The number comes out of your HRIS or IdP record for that identity. Worked when: the field is read-only and the analyst physically cannot type in a number the caller reads out.

3. Hang up. Fully. Not a hold, not a warm transfer — the analyst terminates the call and originates a new one. Worked when: the CDR shows two separate legs with a gap, and the analyst hears ringback, because a held call keeps the attacker's audio path open.

4. Add a channel that is not voice. A push approval in the IdP, or a code delivered to the enrolled device, or a message to the account's own chat profile. Worked when: an approval event appears in the IdP log bound to the enrolled device ID, and you can see it before the reset completes.

5. For the highest-risk actions, require a human who is not the caller. A named manager attests in the ticket. Worked when: the ticket will not close without a second employee ID that differs from the subject's.

6. Give the analyst a no-fault hold code. One phrase — "I need to place this on a verification hold" — that requires no justification and never counts against handle time. Worked when: an analyst uses it on a genuine VP and gets a thank-you from their manager, not a QA note.

7. Log every deviation as a field, not a comment. Worked when: you can run a query for verification-bypassed tickets and get a number for last week. If that number is zero, your logging is broken, not your process.

What each signal actually proves

The caller offers What it proves
A voice that sounds correct Nothing over a phone codec
Employee ID, last four of SSN, manager's name Knowledge that is frequently breached or public
Caller ID matching a corporate range Nothing; trivially spoofed
Convincing office background noise Nothing; may be generated by the network
Push approved from the enrolled device Possession, unless the user was fatigued into tapping
Answering a callback to the directory number Control of the enrolled line — the strongest of these

Rehearse it with a clone you built on purpose

Tabletops do not test whether a 23-year-old will hang up on someone who sounds like the head of trading. Run it live. Get written consent from the executive being impersonated, use their own publicly posted conference audio, clone it with a commercial service under your own account, and call your own help desk from a number outside your ranges. Document the authorization before the first call, not after.

Score one thing: did the analyst complete the callback? Not whether they were suspicious, not whether they felt something was off. Suspicion is not a control. The hang-up is the control.

Where this lands in your detection stack

Mapping it makes the alerting concrete. The pretext is impersonation (T1656). The reconnaissance that supplies the org chart and the voice samples is phishing for information (T1598). The push-fatigue step is MFA request generation (T1621), and the objective is valid accounts (T1078). From that you can write the rule that actually matters: MFA device enrollment within a short window of a help desk ticket, from a device fingerprint unseen on that account, correlated with an inbound call. That correlation is the alert. The audio is not.

About the detection vendors

You will be pitched real-time synthetic-voice detection this quarter. Ask three questions: was the accuracy measured on narrowband, codec-compressed telephony audio or on clean studio files; who ran the evaluation, the vendor or an independent party; and what the false-positive rate does to your queue when a legitimate caller with a cold gets flagged. Most published figures are vendor claims on clean audio, and clean audio is precisely the condition your phone system does not provide. A detector may eventually be worth buying. It will never be worth skipping step three.

Tonight's rule of thumb: if a voice on a call is the reason you are about to change how someone authenticates, hang up and dial the number in the directory — and if that feels rude, that is the control working.

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
J

Juno Park

Game Audio Writer

Juno Park covers AI sound design and game audio workflows — foley, loops, and middleware — after seven years cutting assets for mobile and indie titles. More by Juno Park →