The Illusion of Confirmation
You receive an email from your finance director. It says a confidential acquisition is moving forward, and you will be needed on a secure video call to authorise the initial payment. The email is plausible, referencing real projects and people. You join the call. There is your director, looking and sounding exactly as they always do, perhaps against a slightly blurred background of their home office. They explain the urgency, the need for secrecy, and instruct you to process a substantial transfer to a new account. The video call feels like confirmation. It feels like verification. That feeling is the exploit.
This attack works because it weaponises a natural human heuristic: if I can see and hear someone I trust, I am talking to them. The fraudster uses a pretext—like the email—to set the stage, then uses a deepfake video or a cloned voice to impersonate authority. The request always carries urgency and an injunction to secrecy, which are psychological tools to bypass rational checks. The goal is to make the video call the centre of your reality, making any other procedure feel like an unnecessary delay. The problem is not that the deepfake is perfect; it is that a good-enough fake, deployed in the right context, is perfectly convincing.
Why Spot-the-Fake is a Failed Defence
Many organisations respond to this threat by trying to train staff to detect deepfakes. They run workshops showing slightly off lip-sync, strange eye reflections, or unnatural hair movement. This approach is fundamentally flawed for operational security. You are asking a finance clerk, under pressure from an apparent senior executive, to become a forensic media analyst in real time. The cognitive load is impossible, and the stakes of a false negative—accusing your boss of being a fake—are socially untenable.
Furthermore, the technology is a moving target. The artefacts you learn to spot today are often solved in the next model iteration. Attackers can now use deepfake-as-a-service fraud tools that require no technical skill, offering high-quality outputs for a fee. Relying on human detection shifts the entire burden of defence onto the moment of greatest psychological pressure. It creates a checklist for failure. Your defence cannot be "notice the glitch." It must be "the glitch does not matter, because we never authorise this way."
The Principle: Assume All Biometrics are Compromised
The foundational principle for defeating this fraud is to assume that any biometric signal—face, voice, even a mannerism—can and will be faked. You must design your payment controls with this as a first principle. The video call is not evidence; it is merely the request channel. Verification must happen through a separate, trusted channel that is independent of the request. This is often called an out-of-band process.
This means the identity of the person on the call is irrelevant. What matters is proving the legitimacy of the instruction through a pre-agreed, compromise-resistant method. The goal is to break the attacker's kill chain by inserting a step they cannot intercept or simulate. When you adopt this mindset, you stop worrying about the quality of the fake and start enforcing the strength of your procedure.
Controls That Actually Work
Effective controls are procedural gates that are triggered by the nature of the request, not by your suspicion. They should be simple, mandatory, and culturally reinforced.
- Mandatory Call-Back on a Verified Number: Any payment or account change instruction received via call, email, or message must be confirmed by the authorised staff member placing a call to the requester. The critical rule: the callback number must come from a separate, internal company directory or a previously verified contact list—never from the same email or call that made the request. This defeats most phishing and impersonation.
- Two-Person Approval for Threshold Amounts: Establish clear financial thresholds that automatically require a second, independent authorisation. The second person must perform their own verification, not simply rubber-stamp the first. This control is about requiring collusion from attackers, making the scam far more difficult to execute.
- No Payment Detail Changes via Request Channel: Establish an iron rule: new beneficiary account details are never accepted over the phone, email, or video call alone. They must be submitted through a secure, authenticated portal or via a signed, physical document following a pre-established process. This directly blocks the attacker's goal.
- Pre-Arranged Code Words for Urgent Bypasses: If you must have a process for emergency, out-of-policy payments, it should involve a pre-shared, one-time code word or phrase. The code is established in person or via a highly secure method and is never used over the primary request channel. The person receiving the request must hear the code before proceeding.
- Cultivate a Culture of Expected Verification: Leadership must actively and repeatedly state that they expect to be checked. They must praise, not punish, staff who follow verification procedures, even when it causes delay. This removes the social fear of "bothering the boss" that attackers prey upon.
These controls work in concert. A call-back verifies the requester. A two-person rule provides oversight. Blocking detail changes via email stops the payload. This multi-layered approach is how you build resilience, much like the layered defences needed against help desk social engineering attacks that target IT staff.
A Walk-Through: With and Without Controls
Picture a typical accounts payable officer, Sam. An email arrives from "CFO Anna Lee" about a confidential vendor payment. A video call follows. Deepfake Anna gives urgent instructions to wire funds to a new account.
Without Controls: Sam, seeing the CFO, feels the pressure of urgency. The procedure is vague: "get approval." The video feels like approval. Sam processes the payment. The money is gone.
With Controls: Sam hears the request. The procedure is clear. Any new account detail requires secure form submission; this is a violation. Any payment over the threshold needs a second approver. Sam tells the person on the call, "I need to follow procedure. I will initiate a callback to the number on the internal leadership page to confirm." The attacker cannot control that separate channel. The scam collapses. Sam is not detecting a deepfake; Sam is following a rule that makes the deepfake irrelevant.
This same principle of separating the request from the verification channel is vital for personal security too, which is why understanding how to tell if a voice call is an AI clone relies on similar out-of-band checks. The threat extends beyond finance; consider the implications for deepfake job candidates in remote hiring, where the verification principle must also be rigorously applied.
Questions people ask
Can't we just ban video calls for payment authorisation?
That is a strong operational policy, and many organisations are adopting it. However, attackers are adaptable. A ban on video may simply shift them to using highly convincing voice clones, or to more elaborate email compromise schemes. The more strong solution is to ban the acceptance of instructions via any single, unverified channel. The medium of the request is less important than the rule that the request must be validated elsewhere.
What if the deepfake uses a live, interactive feed?
Some advanced attacks may use real-time puppeteering of a deepfake avatar, allowing for limited interaction. This does not change the defence. Your controls should not be trying to win a "Turing test" with the avatar. Your procedure should be invoking a separate verification step—like the mandatory callback—that the live fake cannot participate in because it does not control the verified phone line.
Do small businesses need these controls?
They are arguably more critical for small businesses, which often have less formal procedures and closer working relationships that attackers exploit. The controls can be scaled: a two-person approval could be the owner and a bookkeeper; a call-back rule is simple to implement. The consequence of a single successful fraud can be existential for a small firm.
Who is ultimately responsible for authorising the payment?
Ultimately, legal and regulatory responsibility lies with the organisation and its authorised signatories. If an employee is socially engineered into bypassing controls, the liability may still rest with the company for having inadequate procedures. This is why documented, enforced controls are not just a technical measure but a governance and liability one.
Close
The deepfake video call scam is not a story about technology outrunning us. It is a story about procedure failing to account for human psychology and technological capability. When you make a video call the centrepiece of trust, you have already lost. The defence is architectural: you build gates in your process that cannot be opened by a face or a voice alone.
Your payment authorisation system must be designed for a world where your colleague's likeness can be rented by a criminal for an hour. The solution is elegantly simple. You decide, in advance, that certain actions require a key that a deepfake can never possess—a connection over a pre-verified channel, a second set of eyes, a code from a different world. You verify, independently, every time. That is the payment nobody should have made, and with the right controls, the one that never will be.
