Towards More General Video-based Deepfake Detection through Facial Feature Guided Adaptation for Foundation Model
With the rise of deep learning, generative models have enabled the creation of highly realistic synthetic images, presenting challenges due to their potential misuse. While research in Deepfake detection has grown rapidly in response, many detection methods struggle with unseen Deepfakes generated by new synthesis techniques. To address this generalisation challenge, we propose a novel Deepfake detection approach by adapting the Foundation Models with rich information encoded inside, specifically using the image encoder from CLIP which has demonstrated strong zero-shot capability for downstream tasks. Inspired by the recent advances of parameter efficient fine-tuning, we propose a novel side-network-based decoder to extract spatial and temporal cues from the given video clip, with the promotion of the Facial Component Guidance (FCG) to encourage the spatial feature to include features of key facial parts for more robust and general Deepfake detection. Through extensive cross-dataset evaluations, our approach exhibits superior effectiveness in identifying unseen Deepfake samples, achieving notable performance improvement even with limited training samples and manipulation types. Our model secures an average performance enhancement of 0.9\% AUROC in cross-dataset assessments comparing with state-of-the-art methods, especially a significant lead of achieving 4.4\% improvement on the challenging DFDC dataset.
Code (1)
Tasks
DecoderDeepFake DetectionFace Swappingparameter-efficient fine-tuningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
One Detector to Rule Them All: Towards a General Deepfake Attack Detection Framework
Deep learning-based video manipulation methods have become widely accessible to the masses. With little to no effort, people can quickly learn how to generate deepfake (DF) videos. While deep learning-based detection met…
AllDeep LearningFace SwappingDetecting Deepfake by Creating Spatio-Temporal Regularity Disruption
Despite encouraging progress in deepfake detection, generalization to unseen forgery types remains a significant challenge due to the limited forgery clues explored during training. In contrast, we notice a common phenom…
DeepFake DetectionFace SwappingA Convolutional LSTM based Residual Network for Deepfake Video Detection
In recent years, deep learning-based video manipulation methods have become widely accessible to masses. With little to no effort, people can easily learn how to generate deepfake videos with only a few victims or target…
DeepFake DetectionFace SwappingTransfer LearningJoint Audio-Visual Attention with Contrastive Learning for More General Deepfake Detection
With the continuous advancement of deepfake technology, there has been a surge in the creation of realistic fake videos. Unfortunately, the malicious utilization of deepfake poses a significant threat to societal moralit…
Contrastive LearningDeepFake DetectionFace SwappingHuman Detection of DeepfakesDeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms
As the GAN-based face image and video generation techniques, widely known as DeepFakes, have become more and more matured and realistic, there comes a pressing and urgent demand for effective DeepFakes detectors. Motivat…
DeepFake DetectionFace SwappingPhotoplethysmography (PPG)Video Generation