paper-with-me

Papers

Joint Audio-Visual Attention with Contrastive Learning for More General Deepfake Detection

2024-01-22 · ACM Transactions on Multimedia Computing, Communications, and Applications 2024 1 · Yibo Zhang, WEIGUO LIN, andJUNFENG XU

With the continuous advancement of deepfake technology, there has been a surge in the creation of realistic fake videos. Unfortunately, the malicious utilization of deepfake poses a significant threat to societal morality and political security. Therefore, numerous researchers have proposed various deepfake detection methods. However, traditional deepfake approaches tend to focus on specific forgery features, such as artifacts or inconsistent actions, which can be vulnerable to specialized countermeasures. Recent studies show an intrinsic correlation between facial and audio cues, which can be exploited for deepfake detection. To address these challenges and enhance the robustness and generalization of deepfake detection algorithms, we propose a novel joint audio-visual deepfake detection model named AVA-CL, which is capable of detecting deepfakes in both audio and visual domains. Furthermore, exploiting the inherent correlation and consistency between audio and visual enhances the effectiveness of deepfake detection significantly. Through extensive experiments, we demonstrate that our proposed AVA-CL model outperforms many state-of-the-art (SOTA) methods with superior robustness and generalization capabilities. This research presents a promising approach for deepfake detection and reducing the harm caused by malicious use.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDeepFake DetectionFace SwappingHuman Detection of Deepfakes

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Contrastive Audio-Visual Masked Autoencoder

2022-10-02 · Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath 외

In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by co…

Audio ClassificationAudio TaggingContrastive LearningMulti-modal Classification+4

Self-supervised Contrastive Learning for Audio-Visual Action Recognition

2022-04-28 · Yang Liu, Ying Tan, Haoyuan Lan

The underlying correlation between audio and visual modalities can be utilized to learn supervised information for unlabeled videos. In this paper, we propose an end-to-end self-supervised framework named Audio-Visual Co…

Action RecognitionContrastive LearningSelf-Supervised Action Recognition

CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling

2023-12-08 · Ruihan Yang, Hannes Gamper, Sebastian Braun

We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditor…

Audio Generation

Learning Self-Supervised Audio-Visual Representations for Sound Recommendations

2024-12-10 · Sudha Krishnamurthy

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of…

Contrastive Learning

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

2025-04-21 · CVPR 2025 1 · Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim 외

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embed…

Contrastive LearningCross-Modal RetrievalSound Source Localization