paper-with-me

Papers

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

2026-05-19 · Aritra Marik, Marcel Klemt, Anna Rohrbach arxiv

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake detection model. Thus, generalizing to deepfakes unseen during training is one of the major challenges in current deepfake detection research. To tackle this challenge, we employ high-level semantic cues and argue that these cues can support low-level focused approaches in generalizing to unseen types of manipulations. In this work, we study emotions as a high-level semantic cue. We propose Emo-Boost, a multimodal deepfake detection framework that fuses an off-the-shelf RGB- and acoustic-focused deepfake detector with our emotion-based deepfake detector EmoForensics. EmoForensics utilises vision and audio emotion recognition modules and models intra- and inter-modal temporal consistency in emotion representations from an audio-visual stream. We found that EmoForensics and the low-level focused method capture complementary signals. Consequently, combining both signals in EmoBoost enhances the average cross-manipulation generalization AUC by 2.1% on FakeAVCeleb.

📄 PDF Abstract BibTeX arXiv:2605.19630

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionDeepFake Detection

Similar Papers 제목 키워드 기반

FauForensics: Boosting Audio-Visual Deepfake Detection with Facial Action Units

2025-05-13 · Jian Wang, Baoyuan Wu, Li Liu, Qingshan Liu

The rapid evolution of generative AI has increased the threat of realistic audio-visual deepfakes, demanding robust detection methods. Existing solutions primarily address unimodal (audio or visual) forgeries but struggl…

DeepFake DetectionFace Swapping

Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition

2024-06-01 · Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian Mishara

Speech emotion recognition (SER) has gained significant attention due to its several application fields, such as mental health, education, and human-computer interaction. However, the accuracy of SER systems is hindered …

Emotion Recognitionfeature selectionSpeech Emotion Recognition

Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content

2024-12-12 · Sheng Wu, Xiaobao Wang, Longbiao Wang, Dongxiao He 외

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within …

Multimodal Sentiment AnalysisSentiment Analysis

Enhancing Infant Crying Detection with Gradient Boosting for Improved Emotional and Mental Health Diagnostics

2024-10-11 · Kyunghun Lee, Lauren M. Henry, Eleanor Hansen, Elizabeth Tandilashvili 외

Infant crying can serve as a crucial indicator of various physiological and emotional states. This paper introduces a comprehensive approach detecting infant cries within audio data. We integrate Wav2Vec with traditional…

Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation

2024-03-18 · Jun Yu, Wangyuan Zhu, Jichao Zhu

In this paper, we present the solution to the Emotional Mimicry Intensity (EMI) Estimation challenge, which is part of 6th Affective Behavior Analysis in-the-wild (ABAW) Competition.The EMI Estimation challenge task aims…