An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection
DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By only changing lip shape to match the given speech, the facial features of identity is hard to be discriminated in such fake talking face videos. Together with the lack of attention on audio stream as the prior knowledge, the detection failure of fake talking face generation also becomes inevitable. Inspired by the decision-making mechanism of human multisensory perception system, which enables the auditory information to enhance post-sensory visual evidence for informed decisions output, in this study, a fake talking face detection framework FTFDNet is proposed by incorporating audio and visual representation to achieve more accurate fake talking face videos detection. Furthermore, an audio-visual attention mechanism (AVAM) is proposed to discover more informative features, which can be seamlessly integrated into any audio-visual CNN architectures by modularization. With the additional AVAM, the proposed FTFDNet is able to achieve a better detection performance on the established dataset (FTFDD). The evaluation of the proposed work has shown an excellent performance on the detection of fake talking face videos, which is able to arrive at a detection rate above 97%.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingFace DetectionFace GenerationFace SwappingTalking Face GenerationSimilar Papers 제목 키워드 기반
FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction
DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By on…
Face DetectionFace GenerationFace SwappingOptical Flow Estimation+1KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of det…
Self-Supervised LearningDeepFake DetectionFrom Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In sing…
DeepFake DetectionOmni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations…
DeepFake DetectionGLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection
Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the u…
DeepFake DetectionFace GenerationFace SwappingTalking Face Generation