paper-with-me

홈 › Papers

BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

2025-05-19 · Haiquan Wen, Yiwei He, Zhenglin Huang, Tianxiao Li, Zihan Yu, Xingru Huang, Lu Qi, Baoyuan Wu, Xiangtai Li, Guangliang Cheng

Advances in AI generative models facilitate super-realistic video synthesis, amplifying misinformation risks via social media and eroding trust in digital content. Several research works have explored new deepfake detection methods on AI-generated images to alleviate these risks. However, with the fast development of video generation models, such as Sora and WanX, there is currently a lack of large-scale, high-quality AI-generated video datasets for forgery detection. In addition, existing detection approaches predominantly treat the task as binary classification, lacking explainability in model decision-making and failing to provide actionable insights or guidance for the public. To address these challenges, we propose \textbf{GenBuster-200K}, a large-scale AI-generated video dataset featuring 200K high-resolution video clips, diverse latest generative techniques, and real-world scenes. We further introduce \textbf{BusterX}, a novel AI-generated video detection and explanation framework leveraging multimodal large language model (MLLM) and reinforcement learning for authenticity determination and explainable rationale. To our knowledge, GenBuster-200K is the {\it \textbf{first}} large-scale, high-quality AI-generated video dataset that incorporates the latest generative techniques for real-world scenarios. BusterX is the {\it \textbf{first}} framework to integrate MLLM with reinforcement learning for explainable AI-generated video detection. Extensive comparisons with state-of-the-art methods and ablation studies validate the effectiveness and generalizability of BusterX. The code, models, and datasets will be released.

📄 PDF Abstract BibTeX arXiv:2505.12620

Code (1)

l8cv/busterx 공식 구현

Tasks

Binary ClassificationDeepFake DetectionFace SwappingLarge Language ModelMisinformationMultimodal Large Language ModelVideo Generation

Similar Papers 제목 키워드 기반

BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM

2025-07-19 · Haiquan Wen, Tianxiao Li, Zhenglin Huang, Yiwei He 외 arxiv

The rapid advancement of generative AI has substantially improved image and video synthesis, amplifying the risk of multimodal visual misinformation. Recent MLLMs have shown promise for transparent AI-generated content d…

Visual Reasoning

Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning

2025-11-22 · Xiaohong Liu, Xiufeng Song, Huayu Zheng, Lei Bai 외 arxiv

The proliferation of videos generated by diffusion models has raised increasing concerns about information security, highlighting the urgent need for reliable detection of synthetic media. Existing methods primarily focu…

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

2026-08-07 · Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou 외 hf

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmark…

Face Swapping

CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

2026-05-26 · Huidong Feng, Wentao Chen, Jie Chen, Xinqi Cai 외 arxiv

With the rapid advancement of artificial intelligence generated content (AIGC) technologies, video forgery has become increasingly prevalent, posing new challenges to public discourse and societal security. Despite remar…

Contrastive LearningDeepFake DetectionVideo Generation

AvatarShield: Visual Reinforcement Learning for Human-Centric Video Forgery Detection

2025-05-21 · Zhipei Xu, Xuanyu Zhang, Xing Zhou, Jian Zhang

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, particularly in video generation, has led to unprecedented creative capabilities but also increased threats to information integrity…

reinforcement-learningReinforcement Learningtext annotationVideo Forensics+1