paper-with-me

Papers

MIS-AVoiDD: Modality Invariant and Specific Representation for Audio-Visual Deepfake Detection

2023-10-03 · Vinaya Sree Katamneni, Ajita Rattani

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has emerged with either audio or visual modalities manipulated. In this regard, a new generation of multimodal audio-visual deepfake detectors is being investigated to collectively focus on audio and visual data for multimodal manipulation detection. Existing multimodal (audio-visual) deepfake detectors are often based on the fusion of the audio and visual streams from the video. Existing studies suggest that these multimodal detectors often obtain equivalent performances with unimodal audio and visual deepfake detectors. We conjecture that the heterogeneous nature of the audio and visual signals creates distributional modality gaps and poses a significant challenge to effective fusion and efficient performance. In this paper, we tackle the problem at the representation level to aid the fusion of audio and visual streams for multimodal deepfake detection. Specifically, we propose the joint use of modality (audio and visual) invariant and specific representations. This ensures that the common patterns and patterns specific to each modality representing pristine or fake content are preserved and fused for multimodal deepfake manipulation detection. Our experimental results on FakeAVCeleb and KoDF audio-visual deepfake datasets suggest the enhanced accuracy of our proposed method over SOTA unimodal and multimodal audio-visual deepfake detectors by $17.8$% and $18.4$%, respectively. Thus, obtaining state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2310.02234

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionFace Swapping

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning

2022-12-10 · Chen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou 외

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant re…

Audio-Visual Speech Recognitionreinforcement-learningReinforcement Learning (RL)speech-recognition+2

The ReprGesture entry to the GENEA Challenge 2022

2022-08-25 · Sicheng Yang, Zhiyong Wu, Minglei Li, Mengchen Zhao 외

This paper describes the ReprGesture entry to the Generation and Evaluation of Non-verbal Behaviour for Embodied Agents (GENEA) challenge 2022. The GENEA challenge provides the processed datasets and performs crowdsource…

DecoderGesture GenerationRepresentation LearningRhythm

Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis

2025-01-07 · Xincheng Wang, Liejun Wang, Yinfeng Yu, Xinxin Jiao

Multimodal Sentiment Analysis (MSA) integrates diverse modalities(text, audio, and video) to comprehensively analyze and understand individuals' emotional states. However, the real-world prevalence of incomplete data pos…

Multimodal Sentiment AnalysisRepresentation LearningSentiment Analysis

Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting

2025-05-22 · Youngmoon Jung, Yong-Hyeok Lee, Myunghun Jung, Jaeyoung Roh 외

For text enrollment-based open-vocabulary keyword spotting (KWS), acoustic and text embeddings are typically compared at either the phoneme or utterance level. To facilitate this, we optimize acoustic and text encoders u…

Keyword SpottingMetric Learning

MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition

2024-01-07 · He Wang, Pengcheng Guo, Pan Zhou, Lei Xie

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learning+3