paper-with-me

Papers

Leveraging large multimodal models for audio-video deepfake detection: a pilot study

2026-02-25 · Songjun Cao, Yuqi Li, Yunpeng Luo, Jianjun Yin, Long Ma arxiv

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale poorly and generalize weakly across domains. We introduce AV-LMMDetect, a supervised fine-tuned (SFT) large multimodal model that casts AVD as a prompted yes/no classification - "Is this video real or fake?". Built on Qwen 2.5 Omni, it jointly analyzes audio and visual streams for deepfake detection and is trained in two stages: lightweight LoRA alignment followed by audio-visual encoder full fine-tuning. On FakeAVCeleb and Mavos-DD, AV-LMMDetect matches or surpasses prior methods and sets a new state of the art on Mavos-DD datasets.

📄 PDF Abstract BibTeX arXiv:2602.23393

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake Detection

Similar Papers 제목 키워드 기반

FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

2021-08-11 · Hasam Khalid, Shahroz Tariq, Minha Kim, Simon S. Woo

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be us…

DeepFake DetectionFace Swapping

Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors

2021-09-07 · Hasam Khalid, Minha Kim, Shahroz Tariq, Simon S. Woo

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. …

DeepFake DetectionFace Swapping

AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Video Deepfake Detection

2023-11-05 · Sahibzada Adil Shahzad, Ammarah Hashmi, Yan-Tsung Peng, Yu Tsao 외

Multimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely d…

DeepFake DetectionFace SwappingSelf-Supervised LearningVideo Forensics

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

2025-08-24 · Xin Zhang, Jiaming Chu, Jian Zhao, Yuchu Jiang 외 arxiv

Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge…

DeepFake Detection

MIS-AVoiDD: Modality Invariant and Specific Representation for Audio-Visual Deepfake Detection

2023-10-03 · Vinaya Sree Katamneni, Ajita Rattani

Deepfakes are synthetic media generated using deep generative algorithms and have posed a severe societal and political threat. Apart from facial manipulation and synthetic voice, recently, a novel kind of deepfakes has …

DeepFake DetectionFace Swapping