paper-with-me

홈 › Papers

AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies

2024-03-22 · Rui Wang, Dengpan Ye, Long Tang, Yunming Zhang, Jiacheng Deng

With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this paper, we propose AVT2-DWF, the Audio-Visual dual Transformers grounded in Dynamic Weight Fusion, which aims to amplify both intra- and cross-modal forgery cues, thereby enhancing detection capabilities. AVT2-DWF adopts a dual-stage approach to capture both spatial characteristics and temporal dynamics of facial expressions. This is achieved through a face transformer with an n-frame-wise tokenization strategy encoder and an audio transformer encoder. Subsequently, it uses multi-modal conversion with dynamic weight fusion to address the challenge of heterogeneous information fusion between audio and visual modalities. Experiments on DeepfakeTIMIT, FakeAVCeleb, and DFDC datasets indicate that AVT2-DWF achieves state-of-the-art performance intra- and cross-dataset Deepfake detection. Code is available at https://github.com/raining-dev/AVT2-DWF.

📄 PDF Abstract BibTeX arXiv:2403.14974

Code (1)

raining-dev/avt2-dwf 공식 구현 pytorch

Tasks

DeepFake DetectionFace Swapping

Similar Papers 제목 키워드 기반

MIS-AVoiDD: Modality Invariant and Specific Representation for Audio-Visual Deepfake Detection

2023-10-03 · Vinaya Sree Katamneni, Ajita Rattani

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has …

DeepFake DetectionFace Swapping

Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization

2024-08-02 · Vinaya Sree Katamneni, Ajita Rattani

In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic …

DeepFake DetectionFace Swapping

AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection

2024-06-05 · CVPR 2024 1 · Trevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 외

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the disso…

Contrastive LearningDeepFake DetectionFace SwappingRepresentation Learning

Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

2025-06-09 · Kuiyuan Zhang, Wenjie Pei, Rushi Lan, Yifang Guo 외

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…

audio-visual learningDeepFake DetectionFace SwappingMisinformation+1

LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization

2026-04-27 · Bokang Zeng, Zheng Gao, Xiaoyu Li, Xiaoyan Feng 외 arxiv

Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often decouple audio and visual evidence and assume that watermark signals…

DeepFake Detection