Question

How are deepfakes made, and can they be detected?

Vault Verified
Curated Intelligence
Definitive Source
Answer

By training a model to map one person's appearance or voice onto another's performance — and detection is possible but increasingly unreliable, which is why the response has shifted toward provenance instead.

How they are produced:

Face swapping, historically using an autoencoder trained on images of two people, learning to reconstruct each face from a shared compressed representation, then decoding one person's expressions with the other's appearance.

Generative models — GANs and, now more commonly, diffusion models — which generate faces and whole scenes rather than transplanting them, and no longer require extensive footage of the target.

Voice cloning, which now requires remarkably little source audio — seconds in some systems — and is the form causing most immediate real-world harm.

Lip synchronisation, altering mouth movement to match new audio, which is subtler and harder to notice than a full face swap.

What detection looks for: inconsistencies in blinking, lighting and shadows; unnatural head and neck boundaries; irregularities where hair meets background; temporal inconsistency between frames; compression artefacts that differ between regions; and physiological signals such as subtle colour changes from blood flow that generators historically failed to reproduce.

Why detection is losing:

Every published detector becomes a training objective, and generators improve against it.

False positives on genuine content are damaging, and compression and re-encoding degrade the signals detectors rely on.

Detection produces a probability, not proof, which is weak evidence in a dispute.

The liar's dividend is the underappreciated harm: once convincing fakes are known to exist, genuine recordings can be dismissed as fabricated. The damage is not only that false things are believed but that true things can be denied.

The responses that scale better: cryptographic provenance attached at capture; watermarking generated output; platform policies requiring disclosure; and — most practically — verification by context, checking whether a recording is corroborated by other sources, rather than examining the pixels.

The legal position has developed, with offences covering non-consensual intimate images and, in some jurisdictions, election-related synthetic media.

Related Questions