paper-with-me

Papers

AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

2025-11-26 · Shuhan Xia, Peipei Li, Xuannan Liu, Dongsen Zhang, Xinyu Guo, Zekun Li arxiv

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and single-granularity annotations, thus failing to capture the diversity and complexity of real-world forgery scenarios. To address this, we introduce AVFakeBench, the first comprehensive audio-video forgery detection benchmark that spans rich forgery semantics across both human subject and general subject. AVFakeBench comprises 12K carefully curated audio-video questions, covering seven forgery types and four levels of annotations. To ensure high-quality and diverse forgeries, we propose a multi-stage hybrid forgery framework that integrates proprietary models for task planning with expert generative models for precise manipulation. The benchmark establishes a multi-task evaluation framework covering binary judgment, forgery types classification, forgery detail selection, and explanatory reasoning. We evaluate 11 Audio-Video Large Language Models (AV-LMMs) and 2 prevalent detection methods on AVFakeBench, demonstrating the potential of AV-LMMs as emerging forgery detectors while revealing their notable weaknesses in fine-grained perception and reasoning.

📄 PDF Abstract BibTeX arXiv:2511.21251

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal Forgery Detection Using Ensemble Learning

2022-11-07 · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) 2022 11 · Ammarah Hashmi, Sahibzada Adil Shahzad, Wasim Ahmad, Chia Wen Lin 외

The recent rapid revolution in Artificial Intelligence (AI) technology has enabled the creation of hyper-realistic deepfakes, and detecting deepfake videos (also known as AIsynthesized videos) has become a critical task.…

Ensemble LearningFace SwappingMultimodal Forgery Detection

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

2025-11-29 · Mengxue Hu, Yunfeng Diao, Changtao Miao, Tairui Ge 외 arxiv

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the v…

Do You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal Method for Temporal Forgery Localization

2022-04-13 · Zhixi Cai, Kalin Stefanov, Abhinav Dhall, Munawar Hayat

Due to its high societal impact, deepfake detection is getting active attention in the computer vision community. Most deepfake detection methods rely on identity, facial attributes, and adversarial perturbation-based sp…

BenchmarkingDeepFake DetectionTemporal Forgery Localization

MSCT: Differential Cross-Modal Attention for Deepfake Detection

2026-04-09 · Fangda Wei, Miao Liu, Yingxue Wang, Jing Wang 외 arxiv

Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily extract forgery traces through audio-visual alignment, which results f…

DeepFake Detection

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

2025-08-13 · Yachao Liang, Min Yu, Gang Li, Jianguo Jiang 외 arxiv

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leverag…

Representation Learning