Security / Synthetic media

Building confidence in deepfake detection

Four years of research, prototyping and independent evaluation helped turn a rapidly changing deepfake threat into repeatable evidence about what detection tools can - and cannot - do.

Back to all case studies
4 yearscontinuous deepfake research and delivery

Problem

Between early 2022 and March 2026, Tekh contributed technical leadership, research and delivery to a long-running public-sector programme focused on synthetic media. The work evolved alongside the technology itself: from early face-swapping and generative adversarial networks to diffusion models and increasingly convincing image, video and audio generation.

The central question remained constant: how can an organisation make defensible decisions about media authenticity when generation techniques change faster than conventional assurance processes?

This was not simply a search for the best detector. We helped build the evidence, methods and reusable infrastructure needed to understand where detection tools worked, where they failed and how their performance changed under realistic conditions.

What we did

Mapping a fast-moving field

Our initial work examined the state of deepfake detection research and the results of earlier technical trials. We reviewed emerging approaches for both still images and video, including spatial and frequency-domain analysis, temporal signals, attention-based methods, model ensembles and techniques designed to generalise across different types of synthetic media.

One early conclusion shaped everything that followed: no single detector performed reliably across every generation method and every kind of media. Results that looked strong on a familiar, clean dataset could deteriorate when the source material, generation technique or processing history changed.

We therefore recommended a more systematic approach built around representative test data, known ground truth and comparable evaluation. Detection needed to be treated as an evolving evidence problem rather than a one-off product selection exercise.

Moving from detection to attribution

As generated images became more realistic, our research expanded from asking whether an image was synthetic to asking how it may have been created.

The idea was analogous to forensic fingerprinting. Cameras and image-generation pipelines leave subtle statistical traces. By analysing images in different domains - including their frequency characteristics and filtered spectra - we explored whether those traces could be represented as fingerprints, grouped in an embedding space and associated with families of generation methods.

We built an early machine-learning prototype combining image transforms, model parsing and classification. In one exploratory evaluation, it was trained on approximately 3,400 images and tested on about 380 images spanning 12 synthetic-image classes. It correctly recognised 97% of the authentic images, with only two synthetic images incorrectly classified as real. These were promising low-readiness results rather than a claim of a solved problem, but they demonstrated that attribution could add useful context beyond a simple real-or-fake score.

We developed the thinking further through a Tekh whitepaper on deepfake image attribution. It argued for combining detection with provenance, model fingerprinting and reverse engineering so investigators and analysts could understand the likely origin of manipulated material and adapt as new generation methods appeared.

Creating fair, independent benchmarks

The next stage focused on turning research into repeatable evaluation. We supported the development and refresh of controlled image, video and audio test data, together with taxonomies describing generation methods, manipulation types, media characteristics and expected ground truth.

This enabled detection tools to be evaluated consistently rather than through vendor-selected demonstrations. The work included:

  • defining representative test scenarios and evaluation criteria;
  • curating real, fully synthetic and partially synthetic media;
  • recording the generation method and relevant parameters for each item;
  • testing tools across different formats, media types and manipulation histories;
  • developing benchmark measures that exposed more than a single headline accuracy score; and
  • presenting results through an interactive leaderboard so performance could be explored by modality, dataset and other characteristics.

Structured feedback showed suppliers where their tools were strong and where additional development or retraining was needed. It also gave decision-makers a clearer basis for comparing capabilities and understanding the operational significance of false positives, false negatives and uncertain results.

Testing beyond clean laboratory data

Real media rarely arrives in pristine form. Images are resized and cropped. Video is compressed, re-encoded or captured from another screen. Audio may be re-recorded. Metadata can be removed, and deliberate post-processing can be used to conceal the traces on which a detector relies.

We incorporated controlled versions of these changes into the test process. This allowed performance to be measured against both straightforward material and evasive or degraded examples. Rather than obscuring the ground truth, every transformation was recorded so that a result could be traced back to the media's source, generation process and processing history.

The wider programme developed a controlled benchmark corpus of around 20,000 assets. Tekh's contribution included scalable generation and orchestration, dataset design, benchmarking support, provenance capture and the creation of repeatable workflows for producing and evaluating multi-modal synthetic media.

Building a refreshable multi-modal capability

By March 2026, the work had moved beyond static datasets. We had developed orchestration capable of generating, annotating and reproducing image, video and audio assets across multiple generation technologies and scenarios.

The approach captured source models, generation settings, transformations and expected ground truth. Automated annotation and validation were combined with human quality assurance, while checksums and structured manifests supported auditability and repeatability. Privacy, ethics and data protection were built into scenario design and production rather than applied as a final check.

This mattered because any fixed benchmark begins ageing as soon as it is created. A refreshable workflow allows new models, manipulation techniques and media formats to be incorporated without rebuilding the entire evaluation process from scratch.

Result

Over four years, the programme moved from horizon scanning to an evidence-led, reusable approach for evaluating synthetic-media detection. Tekh helped establish several principles that remain important:

  • Ground truth is essential. Every result must be connected to a known source and processing history.
  • One score is not enough. Performance needs to be understood by modality, generation method, transformation and use case.
  • Real-world degradation matters. Clean benchmark data alone can create false confidence.
  • Detection and attribution are complementary. Understanding likely origin can add context that a binary result cannot provide.
  • Benchmarks must keep changing. Evaluation capability needs to evolve as quickly as generative technology.
  • Human judgement remains central. Tools should support informed decisions, not replace contextual and forensic interpretation.

The outcome was a stronger, reusable evidence base for comparing deepfake detection capabilities, identifying specific weaknesses and guiding further development. It provided a practical bridge between rapidly advancing research and the careful assurance needed before synthetic-media tools can be trusted in high-consequence environments.

Facing a similar challenge?

Start a conversation with our team.

Start a conversation