paper-with-me

홈 › Papers

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

2026-06-04 · Ruchika Sharma, Rudresh Dwivedi arxiv

Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensive models, which limit their practical use. The study introduces the ExpSpeech-Net deepfake detection (SqN-R-DFD) model, which utilizes SqueezeNet and RNN (Recurrent Neural Network) as its backbone, providing a lightweight and efficient deepfake detection framework that simultaneously analyzes facial expressions and speech patterns. The approach incorporates advanced feature extraction, such as ISLBT-based features for image and MPNCC for signals, along with a smart feature-selection strategy using SASMA (Sandpiper-Assisted Slime Mould Algorithm), ensuring optimal and balanced input to the detection models. By combining SqueezeNet and an RNN, subtle inconsistencies in deepfake videos are captured effectively. The framework achieves 94.5% accuracy, precision of 99.3%, and F-measure of 96.8%, outperforming conventional methods. This demonstrates that integrating multiple modalities with intelligent preprocessing and feature selection enables practical, real-time deepfake detection suitable for everyday applications.

📄 PDF Abstract BibTeX arXiv:2606.05760

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake Detection

Similar Papers 제목 키워드 기반

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

2026-08-24 · Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke 외 arxiv

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigate…

DeepFake DetectionFace Parsing

Diffuse or Confuse: A Diffusion Deepfake Speech Dataset

2024-10-09 · Anton Firc, Kamil Malinka, Petr Hanáček

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We creat…

DeepFake DetectionFace Swapping

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models

2025-06-03 · Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera 외

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they …

Face Swapping

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

2025-05-21 · Yuxuan Du, Zhendong Wang, Yuhao Luo, Caiyong Piao 외

The rapid emergence of multimodal deepfakes (visual and auditory content are manipulated in concert) undermines the reliability of existing detectors that rely solely on modality-specific artifacts or cross-modal inconsi…

cross-modal alignmentDeepFake DetectionFace Swapping

MIS-AVoiDD: Modality Invariant and Specific Representation for Audio-Visual Deepfake Detection

2023-10-03 · Vinaya Sree Katamneni, Ajita Rattani

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has …

DeepFake DetectionFace Swapping