Thesis of Mehdi Atamna
Subject:
Start date: 01/10/2021
End date (estimated): 01/10/2024
Advisor: Serge Miguet
Coadvisor: Iuliia Tkachenko
Summary:
The rapid progress of generative models has enabled the creation of increasingly realistic synthetic facial images and videos, raising significant concerns regarding the authenticity and trustworthiness of digital media. Among these technologies, facial deepfakes represent a particularly challenging form of manipulation due to their potential impact on misinformation, identity fraud, and digital deception. While deep learning-based detection methods have achieved impressive performance under in-distribution evaluation settings, their generalization capabilities remain limited, with significant performance degradation observed when models are evaluated on unseen manipulation techniques, datasets, or generation pipelines.
This thesis investigates how the robustness of facial deepfake detection systems can be improved under challenging generalization scenarios. To address this challenge, this work studies the role of different visual representations, including spatial appearance, temporal dynamics, and frequency-domain information, and investigates how they can be effectively integrated within modern deep learning architectures.
The work conducted in this thesis follows a progressive investigation of deepfake detection robustness. First, the limitations of conventional convolutional neural network-based detectors are analyzed through a cross-dataset evaluation, demonstrating that strong intra-dataset performance does not necessarily translate into reliable performance across different datasets and manipulation methods. Second, the potential of high-frequency information is explored through a lightweight filtering-based approach designed to extract manipulation-related noise residuals. These experiments show that frequency-domain cues provide valuable complementary information, while also highlighting the limitations of relying on a single representation domain.
Building upon these observations, two unified spatio-temporal architectures are introduced. The first proposed approach combines RGB information, wavelet-based frequency representations, and Vision Transformer-based temporal modeling to exploit complementary visual cues within facial video sequences. The second approach further extends this direction through a unified architecture that jointly models spatial and temporal dependencies while also incorporating frequency-domain information through a dedicated wavelet branch. Both architectures are evaluated under challenging cross-dataset scenarios to assess their robustness to unseen manipulations. The second architecture is further evaluated using a standardized benchmarking framework that enables comprehensive comparison with existing approaches, demonstrating the benefits of combining multiple sources of information and achieving competitive performance compared with state-of-the-art methods at a moderate computational cost.