There’s a moment in a well-known 2024 incident that says more about this threat than any statistic could. An employee at Ferrari received a phone call from what sounded, unmistakably, like CEO Benedetto Vigna — right down to his Southern Italian accent — instructing him to handle an urgent, confidential financial matter tied to a supposed acquisition. Everything about the call sounded correct. What stopped the fraud wasn’t a piece of detection software. It was the employee asking one simple, unscripted question: the name of a book the real CEO had recommended to him days earlier. The voice on the other end couldn’t answer, and the call ended immediately. That single moment of procedural instinct is worth more to your security posture right now than any amount of employee training on “how to spot a deepfake.”

Social Engineering Didn’t Get Smarter — It Got Automated

For years, the standard advice for avoiding phishing was to look for the tells: awkward phrasing, generic greetings, a sender domain that was almost right but not quite. Generative AI has quietly erased most of those tells. A model can now draft a flawless, contextually accurate email in the exact tone of an executive, reference real internal projects, and do it in seconds rather than hours. Voice cloning tools need only a short sample of someone’s speech — pulled from a earnings call, a podcast appearance, or a conference keynote — to produce a convincing imitation. None of this requires the attacker to be technically sophisticated anymore. Tools that used to demand real expertise are increasingly available as low-cost, on-demand services, which has meaningfully widened the pool of people capable of running a convincing impersonation campaign.

The result is that the old detection instincts security teams spent years training into employees — check the grammar, listen for a robotic tone, look for unnatural blinking — are becoming unreliable at exactly the moment attackers are scaling their use of these techniques. Independent research on human ability to spot manipulated media has consistently found that people perform only modestly better than chance when asked to identify synthetic audio or video in realistic conditions. That’s not a training gap you can close with a better slideshow. It’s a structural problem with asking humans to be the last line of defense against something engineered specifically to defeat human perception.

Not Every Fake Is a Hollywood-Grade Deepfake

Here’s where a lot of security guidance goes wrong: it treats every synthetic media threat as if it requires a sophisticated GAN pipeline and a technically skilled attacker. In practice, a large share of the damage done by manipulated media comes from something far cruder — what’s often called a “cheapfake.”

A cheapfake doesn’t need generative AI at all. It might be a real video slowed down slightly to make someone sound impaired, a genuine photo shared with a fabricated caption, or old footage re-uploaded and presented as current. These require essentially no technical skill and no budget, yet they’re routinely effective because most people don’t scrutinize context on media they encounter in a moment of urgency or outrage. A company preparing only for high-end synthetic voice clones while ignoring the reputational damage a poorly-edited but widely shared clip can cause is solving half the problem.

The practical takeaway for a security team: your defenses shouldn’t be calibrated only for the most technically impressive attack. The cheap, low-effort version is more common, easier for almost anyone to produce, and still capable of causing real financial or reputational harm.

How These Attacks Actually Get Built

Whether cheap or sophisticated, most of these campaigns follow a similar sequence. It starts with reconnaissance — an attacker gathering publicly available material on a target, often an executive or someone with financial authority. Corporate podcasts, recorded webinars, LinkedIn videos, and conference talks are all fair game, and they’re exactly the kind of content companies actively produce and promote.

From there, the harvested material gets fed into a voice-cloning or face-generation tool to produce the synthetic asset. The attack then moves into delivery, and this is where it gets harder to catch — sophisticated campaigns rarely rely on a single channel. An employee might receive a legitimate-looking email, followed by a voice message on a messaging app, followed by a call, each reinforcing the illusion that the request is real because it’s arriving through multiple, seemingly independent channels.

The final piece is psychological, not technical: manufactured urgency. A request framed as time-critical — a deal closing in the next hour, a confidential matter that can’t wait for normal approval — pushes people to act before their normal skepticism has a chance to engage. This is the same principle that’s driven social engineering for decades; AI has simply made the bait far more convincing.

Two Different Ways Video Gets Faked — And Why It Matters

Within video-based fraud specifically, it’s worth understanding a distinction that security teams researching biometric verification and remote onboarding already know well: presentation attacks versus injection attacks.

A presentation attack happens at the physical level — someone holds a printed photo, a screen replaying a video, or a mask up to a camera, trying to fool whatever is on the other end of that lens. An injection attack skips the camera entirely. Instead, the attacker uses virtual camera software or intercepts the data pipeline to feed a synthetic video stream directly into the system, as if it came from a real camera. Injection attacks are less common but considerably harder to catch, because any defense relying on physical camera characteristics or lighting analysis never gets the chance to evaluate anything real in the first place. Any organization relying on video verification — for identity checks, remote hiring, or KYC onboarding — needs defenses that account for both categories, not just the more visible presentation-style attack.

Building a Defense That Doesn’t Depend on Human Perception

The uncomfortable conclusion from all of this is that training employees to visually or audibly detect a fake is a losing long-term strategy — the technology improves every quarter, and human perception doesn’t. The more durable approach shifts the burden away from perception entirely and onto process.

The single most effective control is a strict, no-exceptions callback policy: any request involving money, credentials, or sensitive access — regardless of how urgent it sounds or how convincing the voice is — gets verified through a separate channel, using a phone number that was saved beforehand, never one provided in the suspicious message itself. This is exactly the instinct that stopped the Ferrari incident, just formalized into policy rather than left to individual employee judgment.

A second, low-cost layer is establishing offline safe words or passphrases for high-stakes transactions — codes that are agreed upon in person or through a trusted channel, never sent digitally, and known only to the people who need them. If a caller claiming to be an executive can’t produce the agreed phrase, the request doesn’t proceed, no matter how real they sound.

Beyond that, requiring dual authorization on financial transfers above a set threshold closes the gap that a single deceived employee would otherwise leave wide open — a second person, on a separate channel, has to independently confirm the request before it’s executed. None of these controls require expensive detection software or a security budget increase. They require deciding, as policy, that no voice or face on a screen is ever treated as sufficient proof of identity on its own.

This isn’t a new idea — it’s the same principle behind zero trust network access applied to human communication instead of network traffic. Assume the call could be fake. Assume the video could be synthetic. Verify through an independent channel every time, and the sophistication of the deepfake stops mattering, because the attack was never going to succeed on convincing appearance alone.