The voice on the phone used to be proof
For most of business and family life, hearing someone's voice or seeing their face on a call was the verification. A written instruction could be forged. A voice could not — not convincingly, not in real time, not without effort far beyond most criminals.
That assumption is now wrong, and it broke quickly. AI voice cloning needs as little as three seconds of audio — a single voicemail, a video posted online, a moment of someone speaking in a public webinar — to produce a convincing replica. Video is only slightly harder: tools that once cost tens of thousands of dollars to build are now free or available for under $20 a month, and a real-time deepfake video call runs on a consumer laptop.
3 seconds
Of audio is enough for current voice-cloning tools to produce a convincing replica of someone's voice.
<30 min
Time needed to clone a voice well enough to fool a trained finance employee, according to fraud researchers.
<$20/mo
Cost of tools that were effectively unavailable to individual criminals five years ago.
What has actually changed is not the scam — impersonating a distressed relative or an urgent executive is one of the oldest tricks in fraud — but the quality bar a criminal now has to clear. It used to take a skilled impressionist. It now takes a public LinkedIn profile.
The $25 million video call
The clearest illustration is also the most expensive one on record. In January 2024, a finance worker at the Hong Kong office of Arup — the engineering firm behind the Sydney Opera House and Beijing's Bird's Nest stadium — received an email from someone claiming to be the company's UK-based CFO, asking for a confidential transaction. He was sceptical, until a video call resolved his doubt: familiar faces, familiar voices, several colleagues he recognised.
Every person on that call was an AI-generated deepfake, built from public footage of real Arup executives at conferences and recorded meetings. Over fifteen transfers to five Hong Kong bank accounts, he sent HK$200 million — about USD 25.6 million. The fraud surfaced only when he later called the company's real head office to follow up. Arup confirmed the incident to CNN in May 2024; the funds were never recovered.
Nobody on the call had been real. Every face he saw and every voice he heard were AI, built from footage anyone could scrape off LinkedIn or a recorded earnings call.
The same pattern runs at household scale, just with smaller numbers and more of them. The FTC's guidance describes it plainly: a call comes in, panicked, sounding exactly like a grandchild or a child — a car accident, an arrest, an urgent need for cash, often with a plea not to tell the rest of the family. The FBI's most recent annual data puts fraud losses reported by Americans over 60 at USD 4.9 billion, up 43% year on year, with impersonation among the fastest-growing categories inside that number.
Lost by Arup to a single deepfaked video call, across fifteen wire transfers.
CNN Business, May 2024Reported to the FBI in business email and impersonation compromise losses in 2024 alone.
FBI IC3 2024 Annual ReportReported fraud losses among Americans over 60 in 2024, up 43% on the year before.
FBI IC3 2024 Annual ReportBusinesses and families are being hit by the same technology aimed at two different vulnerabilities: a company's trust in the chain of command, and a family's instinct to act first and ask questions later when someone they love sounds like they're in danger.
You can't out-look a deepfake
The instinct is to get better at spotting the fake — watch for unnatural blinking, listen for a flat cadence, check whether the lighting matches. That instinct is already running out of runway. The tells that worked in 2023 are mostly gone; the tools have improved faster than most people's ability to notice.
Caller ID is not a defence either. Spoofing a phone number costs a fraction of a cent per call in 2026 and works against every major carrier. A criminal can make the call appear to come from a number already saved in your phone.
Answering a personal question doesn't fully solve it either — not if the question is guessable or has ever appeared online. A pet's name, a childhood street, a school mascot: these sit in data breaches, social media bios, and public records, and a well-resourced attacker can find them faster than most people expect.
What all of this points to is a defence that doesn't depend on detecting the fake at all. It depends on asking for something that was never public in the first place.
A word an AI cannot know
This is the FTC and FBI's actual, current, official guidance: agree a private codeword or short phrase with your family, before you need it. If a call comes in sounding like a relative in distress, the first response is to ask for the word. A cloned voice can say anything — it cannot know something that was never said where a scraper could find it.
Grandparents don't need the whole threat model. They need one instruction: ask for the word, then hang up and call back on the number you already have.
The FBI's guidance on AI-enabled fraud now repeats this almost verbatim, and it appears again in its 2025 alert on scam calls that pose as kidnappings. It is one of the few pieces of fraud advice that requires no technology, no subscription, and no ongoing vigilance — just an agreement made once, while everyone is calm.
The same idea, with a process behind it
A business can't rely on everyone remembering a shared phrase the way a family can — it needs the equivalent built into a process that survives urgency, seniority, and pressure to move fast. Four things do most of the work.
None of this requires new technology. It requires a policy that exists before the call comes in, and staff who have rehearsed following it under the kind of pressure a real attack applies.
The habit that actually holds
Everything here reduces to the same instinct Mono trains for in Click or Flick: stop before you act on urgency, check through a channel the attacker doesn't control, then act. A codeword is that habit made concrete — a single, rehearsed question that works whether the caller is a stranger, a machine, or someone who sounds exactly like family.
It won't stop every attack. Nothing does. But it closes the gap that deepfakes are specifically built to exploit — the moment where a familiar voice replaces a moment's doubt with trust. A five-minute conversation, agreed while everyone is calm, is what stands between a family and a fake emergency call, and between a finance team and a boardroom that never existed.
What this means for your organisation
The tools got cheap fast
Three seconds of audio, freely available tools, and a real-time deepfake call runs on a consumer laptop.
Arup is the ceiling, not the outlier
USD 25.6 million lost to a video call where every participant except the victim was AI-generated.
Detection is a losing race
Visual and audio tells fade faster than most people's ability to spot them. Don't rely on noticing.
Families need one phrase
Agreed in person, never online, never guessable. Ask for it; if it's missing, hang up and call back.
Businesses need a process
Callback verification, dual authorisation, and a codeword for high-risk requests — no exceptions for seniority.
Urgency plus secrecy is the signal
The two levers attackers combine on purpose. A policy that flags both together catches most of the rest.
A codeword only works if the habit is already there.
Click or Flick trains the same instinct this article describes — stop, check, then act — against phishing, deepfakes, and the pressure tactics that make people skip the pause. It measures whether the habit actually held, not just whether the training was completed.