Express Computer
Home  »  Guest Blogs  »  Why detecting deepfakes is structurally harder than building them, and what it will take to close the gap

Why detecting deepfakes is structurally harder than building them, and what it will take to close the gap

0 0

By Ankush Tiwari, Founder and CEO, pi-labs

Somewhere in India this month, a man is watching a video of himself say something he has never said.

The voice is his. In the video he is recommending an investment scheme, and thousands trust him precisely because he has spent years telling them which schemes not to trust. This is roughly what happened to Ankur Warikoo.

In 2026 the Delhi High Court restrained unknown defendants from using his likeness through AI, and ordered Meta to pull the URLs within thirty-six hours. Note what the court had to assemble: no dedicated deepfake statute, so personality rights, trademark law, and a John Doe order against defendants nobody could name. An improvised remedy, built after the damage.

Warikoo is one of the lucky ones, a public profile, a legal team, a court that will hear him. The 47% of Indian adults hit by an AI voice-cloning or deepfake scam, or who know a victim, get none of that. That is nearly double the global average; 83% of Indian victims lost money, almost half more than ₹50,000. If it were your mother’s voice on the phone at 11 PM, asking for help, would you have known?

The global numbers don’t land: eight million deepfake files circulating, contact-centre attempts up 1,300% in a year. So here is a smaller one: 27% of identity-document attacks worldwide involved Indian IDs. More than a quarter of the planet’s document fraud, aimed at one country’s paperwork. 65% of Indian organisations report they have already faced a deepfake-driven attack. And a viral advertisement circulated in which India’s Finance Minister appeared to endorse a fraudulent investment scheme, not an anonymous CEO, but the person whose job is the credibility of Indian financial markets.

India is the hardest deepfake environment on earth, structurally: 120+ languages, and the world’s largest WhatsApp base functioning as an encrypted, unmoderated, infinitely forwardable distribution layer. A cloned voice note in Bhojpuri, forwarded into a family group at midnight, does not meet a content moderator. It meets an uncle who trusts it.

None of this was a failure of foresight. It was written into the technology at its invention. In 2014, Goodfellow’s GAN paper set two networks against each other, one generating, one detecting. Read it closely and something uncomfortable appears: the generator is, by definition, optimised to defeat a detector. Not as a side effect. As the objective function. The arms race isn’t a consequence of the technology. It is the technology.

In 2017, DeepFaceLab reached GitHub, and offence opened a head start it has never given back. In 2022 Stable Diffusion shipped free, and every detector trained on GAN artefacts was suddenly hunting the wrong fingerprints. By 2026, real-time face replacement runs on consumer hardware at latency below human perception, from three seconds of audio. Each advance didn’t merely improve quality.

Each one destroyed the signal detection was using.Why the defender is not playing the same game
“Arms race” implies symmetry. It is the wrong frame. Creation has certainty; detection has probability.

The forger controls the source footage, the model, the compression, the channel. He can generate a thousand variants, test each against every public detector, and ship only the survivor, unlimited attempts, perfect feedback. The detector gets one artefact, no knowledge of what made it, and one attempt. And it must be right in a way the attacker never has to be: crying wolf on authentic evidence destroys trust as thoroughly as missing a fake.

Which is why every published detection signal has an expiry date, and publication starts the clock: the moment a paper documents a reliable artefact, that artefact becomes a training objective for the next model. Detection research functions as free quality assurance for the offense.

Then the failure nobody markets, detectors claiming 90–96% accuracy in the lab drop 45–50% in production. A tool advertising G6% may be delivering 48% in the field. The benchmark is a curated dataset; reality is a 240p WhatsApp forward, twice re-encoded, screen-recorded off another phone, with a regional-language voiceover burned in. A model never stress-tested on an eleven-times-forwarded Bhojpuri clip is not a detection system, it’s a demo. The human fallback sits at 55–60%: a coin flip with extra steps.

One correction my own industry should make: most “AI fraud” in India today is not deepfake fraud.

Digital arrest scams run largely on coercion, rented uniforms, fake police-station sets, industrialised scripts, hours of pressure. Synthetic media appears in some; it is rarely load-bearing. I say this against commercial interest, because the implication is worse: if people surrender their savings to a man in a costume on a laggy call, what happens when the face is synthetic and the latency is gone? Deepfakes don’t create the vulnerability. They industrialise one that already exists.

Leave A Reply

Your email address will not be published.