What Is Deepfake Phishing and How to Spot It

Introduction

A finance employee at the engineering firm Arup joined a routine-looking video call with his chief financial officer and several colleagues, then followed their instructions to wire $25 million. Every person on that call was synthetic. Incidents like this one have moved from research-lab demonstrations to boardroom losses in under three years, and most of the security tools built to catch phishing were never designed to look at a face or listen to a voice. This guide breaks down how these attacks are actually built, the specific things that still give away a fake on video and audio, and the single verification habit that holds up even when the fake doesn’t.

Quick Answer

Deepfake phishing uses AI-generated audio, video, or images to impersonate a trusted person a boss, colleague, or family member and pressure the target into sending money or approving a transaction. Deepfake phishing attacks target what people see and hear, not what they read, so spotting one takes new verification habits, not sharper proofreading.

Key Takeaways

Deepfake phishing slips past spam filters because the payload is a fake voice or face, not a malicious link or code a gateway can scan.

The FTC warns that a small clip of someone’s real speech a voicemail greeting or a short public video is enough to clone their voice convincingly.

The old advice to listen for a robotic tone or watch for stiff blinking no longer reliably works, because current voice and video models have largely fixed both tells.

A verified case at engineering firm Arup shows a deepfake video call can fool trained finance professionals into wiring roughly $25 million, not just distracted individuals.

Caller ID proves nothing, since displaying a trusted name or number on a screen is trivial to fake and doesn’t confirm who is actually speaking.

The one defense that survives a flawless fake is verifying the request on a channel the attacker doesn’t control, such as calling back a known number.

What Is Deepfake Phishing?

Deepfake phishing is a cyberattack that uses artificial intelligence to generate a fake voice, video, or image of someone the target trusts, then uses that fabrication to manipulate the target into an action such as transferring money, sharing a password, or approving a request. It belongs to the broader category of social engineering, manipulating a person rather than a system, and it inherits every classic social engineering trigger, urgency, authority, familiarity, while adding a layer traditional phishing training never covered: audio-visual proof that used to be hard to fake.

Traditional phishing relies on text: a spoofed sender name, a link to a fake login page, or a message with small inconsistencies a trained eye can catch. Deepfake phishing swaps text for sight and sound. Instead of a suspicious email from “the CEO,” the target gets a voice note, a phone call, or a live video feed that sounds and looks like the real person. An AI voice cloning scam and a deepfake video call exploit the same shortcut: people trust their own eyes and ears more than they trust a spam filter.

Deepfake phishing shows up in several formats, and most attacks combine more than one:

  1. Voice calls or voicemails: a cloned voice of a boss, relative, or bank representative, built from a few seconds of audio.
  2. Live video calls: a real-time deepfake face and voice layered onto a video conference, as in the Arup case.
  3. Pre-recorded video: a fabricated clip sent as “proof,” often screen-recorded so the file’s origin can’t be checked.
  4. Fake profiles and messages: a cloned executive’s face on a professional networking profile or messaging app, used to build trust before the ask.

Whichever format arrives first, the goal is the same: replace the target’s skepticism with recognition.

How Deepfake Phishing Attacks Actually Work

A deepfake phishing attack moves through four stages: gathering source material on the target, generating the synthetic voice or video, delivering it through a channel that looks legitimate, and exploiting the trust it built to extract money or data. The third stage, delivery, is exactly where most corporate security tools stop looking.

Each stage exploits a different weakness, and none of them require breaking into a company’s systems:

  1. Data collection: attackers scrape public audio and video: earnings calls, webinars, conference talks, social clips, even a voicemail greeting.
  2. Synthesis: AI models, commonly built on generative adversarial networks and autoencoders, turn that raw material into a cloned voice or a real-time deepfake face.
  3. Delivery: the fake arrives through a trusted-looking channel: a phone call, a video-conference invite, a voicemail, or a message from a familiar-looking account.
  4. Exploitation: the target, convinced by the familiar voice or face, transfers funds, shares credentials, or approves a request they would have questioned from a stranger.

Because nothing in that chain is a malicious link or a piece of code, email gateways and endpoint protection built to catch traditional phishing have nothing to flag which is also why they sit inside the 2025 cybersecurity threat landscape as one of the fastest-growing categories rather than a solved problem.

A Real Deepfake Phishing Case: The $25 Million Arup Call

In January 2024, a finance employee at the Hong Kong office of engineering firm Arup made 15 transfers totalling roughly $25 million after joining a video call with people who looked and sounded exactly like the company’s chief financial officer and several colleagues. None of them were real.

According to CNN’s reporting on the incident, confirmed by an Arup spokesperson, the attack began with a phishing email impersonating the UK-based CFO and requesting a confidential transaction. The employee was initially skeptical of the email, which is exactly what phishing training is supposed to produce, but the doubt dissolved once he joined a video call and saw and heard several familiar colleagues discussing the request. Arup later said its internal systems were never compromised; the attackers didn’t need to breach anything, because the employee let them in.

The same pattern plays out at a smaller scale constantly, even when it never makes headlines: a voicemail that sounds exactly like a department head asking a colleague to buy gift cards for an urgent vendor payment is the same attack with a cheaper production budget. Consistent training employees to recognise social engineering is still the first line of defense, because a deepfake only works if the target never pauses to question it.

How to Spot Deepfake Phishing on a Live Video Call

On a video call, a deepfake usually still leaves seams, but the seams that mattered a few years ago, stiff blinking, a robotic voice, have mostly been fixed by current models. The checks that still work focus on physics and behaviour rather than obvious glitches.

Watch for these signs together, since any single one can also happen on a genuine call:

  1. Lighting and shadow mismatches: the face’s lighting doesn’t shift the way the room’s lighting would.
  2. Edge and texture errors: a blurred hairline, jewelry that flickers or disappears, teeth that look like a single block instead of individual teeth.
  3. Audio-visual drift: lip movements land slightly ahead of or behind the words, especially during fast speech.
  4. Flat or oddly clean audio: studio-quality sound from someone who claims to be outdoors, in a car, or on a bad connection.
  5. Resistance to simple behavioural tests: ask the person to turn their head in profile, hold a hand in front of their face, or say an unscripted word; live deepfakes struggle more with unplanned movement than with scripted speech.
  6. Refusal to switch channels: a sudden excuse when asked to confirm the request by phone or in person instead.

No single sign proves a call is fake, but a caller who fails several of these at once, especially the behavioural ones, has earned a callback on a number you already had.

How to Spot an AI Voice Cloning Scam

An AI voice cloning scam usually arrives as a phone call or voicemail with no video to inspect, so the checks shift from what you see to what the request itself sounds like and asks for.

The voice can be flawless; the script rarely is. Watch for:

  1. Manufactured urgency: a deadline measured in minutes, tied to a payment (“before the bank closes,” “before I miss my flight”).
  2. Enforced secrecy: “don’t tell anyone” or “keep this between us,” which cuts off the exact verification step that would expose the fake.
  3. An unusual payment method: gift cards, cryptocurrency, or a wire to an account the caller has never used before.
  4. A caller ID that matches: spoofing a phone number to display a real contact’s name costs an attacker almost nothing, so a familiar number on the screen proves nothing about who’s speaking.
  5. Refusal or excuses to move to video: “bad connection,” “camera’s broken,” or simply changing the subject when asked.
  6. Emotional pressure that escalates when questioned: a real relative asked to slow down usually can; a cloned voice under attacker control often can’t adapt.

A small clip of someone’s real speech is enough to build a convincing clone. The FTC’s family emergency scam guidance warns that a voicemail greeting or a short public video can give scammers all the raw material they need, which means the source clip could be one the target never thought twice about posting.

The One Defense That Beats a Flawless Fake: Independent Verification

Every defense that matters against deepfake phishing comes down to one habit: verify the request on a channel the attacker doesn’t control, before acting on it, no matter how convincing the call looked or sounded. The two mistakes that show up in almost every reported loss are trusting caller ID as proof of identity, and treating a live video call as automatically more trustworthy than a text message.

For individuals, that means calling the person back on a number saved from before the suspicious contact, never one provided during the call. Families dealing with emergency-call scams can go further by agreeing on a safe phrase in advance, something specific enough that it wouldn’t appear in a scraped social clip and never posted anywhere online. For finance and HR teams, the same idea scales into policy: no wire transfer, credential reset, or “confidential” request gets approved from a single video or voice channel, regardless of who appears to be asking, and every team should already have free antivirus and phishing-protection tools covering the more conventional side of the same threat.

Build the habit before you need it:

  1. Save direct numbers for anyone whose voice could plausibly be cloned to request money; don’t rely on caller ID or a number texted to you in the moment.
  2. Agree on a safe phrase with family or a close team, chosen offline and never posted publicly.
  3. Require a second approval channel for any wire transfer, gift-card purchase, or credential change tied to an urgent request.
  4. Treat “don’t tell anyone” as a red flag on its own, since it’s designed to remove your ability to verify.

None of this requires spotting the fake in the moment it just requires refusing to act until a second, independent channel confirms it.

What to Do If You Suspect a Deepfake Phishing Attempt

If a call, video, or message feels off, stop before sending money, sharing a code, or clicking anything, then work through verification and reporting in order. The sequence matters more than any single step:

  1. End the interaction without confirming or denying your suspicion out loud “let me call you right back” is enough.
  2. Call the person back on a number you already had, not one given to you during the call.
  3. If you can’t reach them directly, verify through a second person a colleague, another family member, or the person’s employer.
  4. Preserve evidence: screenshots, call logs, voicemail recordings, and messages, before they’re deleted or overwritten.
  5. Report it. In the US, that’s the FTC at ReportFraud.ftc.gov, which has also pushed for a proposed ban on impersonation fraud specific to AI-generated calls; most countries have an equivalent national fraud reporting line.
  6. Notify your bank or employer immediately if any money or credentials were already shared, since speed affects whether a transfer can be reversed.

Acting on this sequence in the first hour, not the first day, is usually what determines whether a deepfake phishing attempt costs nothing or costs everything.

Conclusion

Deepfakes will keep getting better at fooling eyes and ears, which is exactly why the defense in this guide doesn’t depend on either. Deepfake phishing attacks succeed only when the target skips verification, so treat every unexpected request for money, credentials, or secrecy, however convincing the voice or face, as a prompt to hang up and call back on a number you already trust. This is the same discipline that stops ordinary social engineering, just applied to a harder-to-spot delivery method. Before your next urgent call catches you off guard, agree on a safe phrase with the people most likely to be impersonated to you, and save their direct numbers somewhere you’ll actually check.

FAQs

1. Can deepfake phishing bypass multi-factor authentication?

Not directly. Deepfake phishing targets a person’s judgment, not a login system. But it often leads to the same outcome: an employee socially engineered by a fake voice or face may reset an MFA device, approve a push notification, or read a one-time code aloud, handing over the exact protection MFA was meant to provide.

2. How much audio does it take to clone someone’s voice?

Very little. The FTC warns that scammers can generate a convincing clone from just a small clip of someone’s real speech a voicemail greeting or a short public video is enough. That’s why security researchers now recommend keeping voicemail greetings generic rather than recorded in your own voice.

3. Are deepfake detection tools reliable?

No detection tool is foolproof. New generative models are trained specifically to defeat the previous generation of detectors, so a “real” score from any single tool is a data point, not proof. Detection software is a useful supplement to verification habits, not a replacement for calling the person back.

4. Is deepfake phishing illegal?

Yes. Deepfake phishing combines fraud and impersonation, both prosecutable under existing laws in most countries, even where no specific “deepfake” statute exists. Regulators including the FTC have also proposed dedicated rules aimed at AI-enabled impersonation fraud and AI-generated scam calls.

5. Can deepfakes fool facial recognition or liveness checks?

Some can. Liveness checks that only ask a user to blink or turn their head have been defeated by real-time deepfakes in documented cases, which is why banks and identity-verification vendors increasingly combine several signals, device, behaviour, and document checks, rather than relying on a face alone.

6. Who do deepfake phishing attacks usually target?

Finance and HR teams are the most common corporate targets, since they can authorise payments or reset credentials, but families are targeted just as often through fake emergency calls impersonating a relative. Anyone with public audio or video online, which is nearly everyone, is potential source material.

logo-white.png

Subscribe to Our Newsletter