AUGUST 18, 2026
At nine in the morning a finance employee joins a routine video call. The chief financial officer is on screen. So are several colleagues from head office. The faces are familiar, the voices are familiar, the small talk before the agenda is ordinary. The instruction that follows is urgent and confidential: a transaction is in progress, the documentation will follow later, the transfers need to leave today. The employee had received an email about this a few days earlier and had been suspicious of it. By the end of the call the suspicion is gone, because everyone in the room confirms the story. Fifteen transfers later, 25.6 million dollars has left the company.
Not one participant on that call was a real person. This is the defining property of a deepfake video call: it does not defeat a technical control, it defeats the human check that every payment process keeps in reserve for exactly this situation. The employee did the right thing by escalating a suspicious email to a live conversation. The live conversation was the attack.
Business email compromise cost organisations 3.04 billion dollars in reported losses in 2025, up from 2.77 billion the year before, and 86 percent of those losses moved by wire or ACH transfer. The email half of that problem has a decade of defensive investment behind it. The half where the forgery arrives as a face on a screen does not.
| $25.6M lost in a single deepfake video call, moved across 15 separate transfers | 62% of organisations faced a deepfake attack in the preceding 12 months | 0.1% of people could reliably tell real content from AI generated content | $3.04B in business email compromise losses reported in 2025 (FBI Internet Crime Report,2025) |
Classic business email compromise was a text problem, and text turned out to be defensible. Authentication protocols made sender spoofing harder. External sender banners taught people to slow down. Lookalike domain monitoring caught the near miss registrations. Finance teams were trained to call back before releasing a payment. None of this eliminated the fraud, but it moved the economics: an attacker had to work harder for the same wire.
So the pretext moved to the channel that carries the most trust per second. A face and a voice, live, responding in real time, are the strongest identity signal most people have ever had access to. Every verification habit built over the last ten years quietly assumed that a video call was the place where doubt goes to be resolved. That assumption is now a control gap.
In January 2024 an employee in the Hong Kong office of the engineering firm Arup received a message about a confidential transaction. The employee suspected phishing, which is exactly the intended outcome of security awareness training. A video conference was then arranged. On it appeared the chief financial officer and several recognisable colleagues. The employee proceeded, and authorised 15 transfers to five bank accounts totalling the equivalent of 25.6 million dollars. Every participant other than the victim was synthetic.
The detail worth dwelling on is not the amount. It is the sequence. The suspicion fired correctly, the escalation path was followed correctly, and the escalation path was the trap. Any organisation whose payment verification ends with ‘confirm it with them directly on a call’ has the same sequence written into its procedure.
Treating this as magic makes it undefendable. It is a supply chain with three distinct stages, and two of them are visible to the target organisation before anything happens.
Synthesis needs source material, and senior executives generate it professionally. Earnings calls, conference keynotes, webinars, recruitment videos, podcast appearances and press interviews produce hours of well lit, high resolution, frontal footage of exactly the people who can authorise large payments. Voice cloning is cheaper still: usable models are built from a few minutes of clean speech, and a single conference panel supplies that.
The second collection target is organisational structure. Who reports to whom, who holds signing authority, which finance platform is in use, which banking partners appear in press releases, who is travelling and therefore plausibly joining from a hotel with poor lighting. Conference speaker lists, job postings that name internal systems, and public profiles supply most of it without a single intrusive action.
The third is timing. Quarter end, an announced acquisition, a known executive absence or a genuine confidential project all provide a pretext that survives casual scrutiny, because the story is partly true.
A live face swap runs as a six stage pipeline: frame capture, face detection, facial landmark extraction, neural synthesis of the target identity in the source pose, blending of the generated face into the frame, and output through a virtual camera driver that presents itself to the conferencing application as an ordinary webcam.
The entire chain has to complete inside roughly 30 to 50 milliseconds per frame to hold a normal 24 to 30 frames per second. A consumer graphics card with 8 to 12 gigabytes of memory is sufficient for real time swapping at 720p, and voice cloning adds a further 200 to 500 milliseconds of end to end latency. This latency budget is the most operationally useful fact in this entire article, because it forces the attacker to run smaller, lower resolution, quantised models, and those compromises are what a trained observer can still see.
The synthetic media is the smallest part of the operation. Around it sits a familiar structure: a calendar invitation from a domain one character away from the real one, a confidentiality frame that explains why normal channels must be avoided, urgency that compresses the time available to verify, and seniority that makes questioning socially expensive.
The one genuinely new element is population. A single fake face asks the victim to trust one person. Five fake faces ask the victim to trust a consensus, and consensus is far harder to argue with. Manufacturing a room full of agreeing colleagues is now a rendering cost rather than a recruitment problem.
Walk a deepfake video call through a typical control stack and it passes every gate, because it is not attacking any of them.
| CONTROL | WHY IT DOES NOT FIRE |
| Email security gateway | There is no email at the decisive moment. The instruction is spoken, and inspection has nothing to parse. |
| Multi factor authentication | MFA proves control of an account. The attacker is not using the executive’s account, they are wearing the executive’s face on their own. |
| Callback verification | It works only when the number comes from an authoritative directory. A number taken from the invitation, the chat panel or the signature block routes back to the attacker. |
| Payment approval workflow | It was followed. A senior person requested it, the amount sat inside authority or was split beneath the threshold, and the approval was logged. The audit trail looks clean afterwards. |
| Security awareness training | Staff were trained to distrust text and to escalate to a live conversation. The live conversation is now the payload. |
A useful way to state the problem: none of these controls verify a human being. They verify accounts, addresses, amounts and approvals. The deepfake attacks the one link everybody assumed was self verifying.

The latency budget described above leaves signatures. These are current weaknesses and they narrow every year, so they belong in your training material but never in your control design. Recognising an artifact is a bonus. The protocol in the next section is the actual control.
| SIGNAL | WHAT TO WATCH FOR |
| Boundary instability | Jitter at the edge of the face, a faint halo around the hairline, colour discontinuity at the neck, background pixels warping during head rotation, ears distorting at sharp angles. |
| Lighting mismatch | The face is lit differently from the room visible behind it. Reflections in the eyes do not match the light sources in the scene, because the model inherited lighting from its training data. |
| Gaze and blink behaviour | Pupils that do not respond to changes in screen brightness, gaze that lags behind head movement, both eyelids closing in perfect synchrony, blink rates that are too regular or too rare. |
| Audio and video drift | Lip movement and sound separated by a fixed offset that never varies. Real speech drifts by tiny amounts; a synthesis pipeline holds a constant delay. |
| Occlusion failure | A hand moving across the face causes flicker, smearing or a momentary collapse of the swap, because landmark detection loses the reference points. |
| Profile limits | Quality degrades sharply when the head turns beyond the angles well represented in training data. |
Each of these exploits a specific constraint in the pipeline rather than the observer’s instinct, which is why they outperform ‘does this look right to you’.
Ask the person to turn their head fully to profile and hold the position while continuing to speak. Ask them to pass a hand slowly across their face. Ask them to stand up and step back so the camera sees more of the room. Ask them to write a word of your choosing on paper and hold it up to the lens. Each request forces the pipeline outside its comfortable operating range in real time.
The strongest challenge is not technical at all. In 2024 an attempted executive impersonation at Ferrari collapsed when the executive receiving the call asked the caller to name the title of a book he had recommended a few days earlier. The call ended immediately. Shared private context cannot be scraped, cannot be rendered, and costs nothing to deploy.
One caution: do not turn these into a script that circulates internally as ‘the deepfake test’. Anything predictable is something the attacker can prepare for. Vary the request.
Artifact spotting is a skill that decays and an advantage that expires. Process does neither. These five controls assume the impersonation is perfect and still prevent the loss.
| CONTROL | WHAT IT PREVENTS |
| 1. Video is never an authorisation channel | State it as policy: no payment instruction, banking detail change or credential action is ever final on a live call, regardless of who is on screen. The call may discuss it. A separate system executes it. This single rule would have stopped the 25.6 million dollar case outright. |
| 2. Out of band callback to a directory number | Verification uses a number retrieved from the corporate directory, never one supplied in the invitation, chat window, email signature or by the caller. The attacker controls every channel they introduced. |
| 3. Dual authorisation with independent initiation | Two approvers who did not attend the same meeting, with a threshold low enough that splitting a transfer into fifteen pieces does not slip beneath it. Fifteen transfers to five accounts should itself be an alerting condition. |
| 4. Beneficiary cooling period | Any new or amended beneficiary is held for a fixed window and confirmed with a known contact at the counterparty using previously established details. Urgency is the attacker’s core asset, so remove the ability to spend it. |
| 5. Rotating challenge phrase for executives | A private phrase per executive, exchanged out of band, never spoken in a recorded meeting, refreshed on a schedule. Cheap, fast, and effective against a perfect visual forgery. |
There is a sixth control and it is cultural rather than technical. These attacks work because a junior employee will not challenge a senior face. Unless the chief executive states plainly, in writing, that verifying an instruction is never a career risk and that no one will ever be penalised for slowing a payment down, the other five controls will be bypassed politely by people who are trying to be helpful.
Synthesis is the final step, and it is the only step defenders cannot influence. Collection is the first step, and it is entirely measurable. The practical question for a security team is not ‘can someone deepfake our CFO’, because the answer is yes. It is ‘how much material exists, how current is it, and what else does an attacker have that makes the pretext credible’.
That reduces to a concrete inventory: how many hours of high quality public video and audio exist for each executive and how recent they are; whether executives’ personal email addresses, phone numbers and home details are exposed in breach data; which domains have been registered against your brand or your executives’ names and which of them resolve to a working conferencing or login page; which fake profiles are impersonating your leadership on professional and messaging platforms; and whether your executives are being discussed by name in closed channels where synthetic media services and cloned voice models are advertised.
Executives are roughly twelve times more likely to be targeted than other employees, a pattern visible in industry breach data to which Brandefense contributes as a data partner. The exposure that makes a deepfake convincing is not a secret held by the attacker. It is public information that nobody on the defending side has counted. (Source: Verizon Data Breach Investigations Report, 2026)
| CAPABILITY | WHAT IT ADDRESSES IN THIS ATTACK CHAIN |
| VIP and Executive Security | Continuous monitoring of the personal digital footprint of named executives: exposed personal accounts, leaked contact details, breach data appearances and impersonation attempts across platforms. |
| Brand Protection and DRPS | Detection of lookalike domains registered against the corporate brand and executive names, including domains built to host convincing meeting invitations and credential capture pages. |
| Dark Web Monitoring | Visibility into closed forums and messaging channels where synthetic media services, cloned voice models and executive targeting packages are traded, including mentions of your organisation by name. |
| Cyber Threat Intelligence | Tracking of the actors and tradecraft behind executive impersonation campaigns, mapped to observable techniques so detection engineering has something concrete to build on. |
| EASM | Identification of forgotten subdomains, expired certificates and orphaned infrastructure that give a fraudulent pretext an authentic looking home. |
RELATED READING
Why Your CISO Is Your Organization’s Highest-Value Attack Target: https://brandefense.io/blog/why-your-ciso-is-your-organizations-highest-value-attack-target/ The exposure inventory this attack depends on, mapped executive by executive.
The Executive Cyber Threat: Personal Email, Home Networks, and Family Members as Attack Vectors: https://brandefense.io/blog/the-executive-cyber-threat/ Where the personal half of an executive footprint turns into a corporate entry point.
VIP Security Beyond the Bodyguard: Why Digital Protection Is the New Executive Priority: https://brandefense.io/blog/executive-digital-protection/ How to run a protection programme that treats digital and physical risk as one problem.
How Spear Phishing Campaigns Target C-Suite Executives: Tactics, Tools, and Defense: https://brandefense.io/blog/spear-phishing-c-suite-executives/ The intelligence gathering stage that makes an impersonation credible before any synthesis starts.
The uncomfortable conclusion is that identity verification by sight and sound, the method humans have relied on for the whole of history, has stopped being sufficient for financial decisions. Organisations that accept this early will rewrite a handful of payment procedures and add a challenge phrase. Organisations that accept it late will do the same work afterwards, with an incident report attached and a number in it.

Take control of your digital security with an exclusive demo of our powerful threat management platform.