paper-with-me

홈 › Papers

Can Multi-modal (reasoning) LLMs work as deepfake detectors?

2025-03-25 · Simiao Ren, Yao Yao, Kidus Zewde, Zisheng Liang, Tsang, Ng, Ning-Yau Cheng, Xiaoou Zhan, Qinzhe Liu, Yifei Chen, Hengwei Xu

Deepfake detection remains a critical challenge in the era of advanced generative models, particularly as synthetic media becomes more sophisticated. In this study, we explore the potential of state of the art multi-modal (reasoning) large language models (LLMs) for deepfake image detection such as (OpenAI O1/4o, Gemini thinking Flash 2, Deepseek Janus, Grok 3, llama 3.2, Qwen 2/2.5 VL, Mistral Pixtral, Claude 3.5/3.7 sonnet) . We benchmark 12 latest multi-modal LLMs against traditional deepfake detection methods across multiple datasets, including recently published real-world deepfake imagery. To enhance performance, we employ prompt tuning and conduct an in-depth analysis of the models' reasoning pathways to identify key contributing factors in their decision-making process. Our findings indicate that best multi-modal LLMs achieve competitive performance with promising generalization ability with zero shot, even surpass traditional deepfake detection pipelines in out-of-distribution datasets while the rest of the LLM families performs extremely disappointing with some worse than random guess. Furthermore, we found newer model version and reasoning capabilities does not contribute to performance in such niche tasks of deepfake detection while model size do help in some cases. This study highlights the potential of integrating multi-modal reasoning in future deepfake detection frameworks and provides insights into model interpretability for robustness in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2503.20084

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionFace Swapping

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection

2026-01-02 · Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra, Abhijit Das 외 arxiv

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored.…

Audio Deepfake Detection

PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection

2025-09-30 · Tuan Nguyen, Naseem Khan, Khang Tran, NhatHai Phan 외 arxiv

The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the scarcity of large, high-quality datasets. Although multimodal large lang…

Reinforcement LearningMultimodal ReasoningDeepFake Detection

Can ChatGPT Detect DeepFakes? A Study of Using Multimodal Large Language Models for Media Forensics

2024-03-21 · Shan Jia, Reilin Lyu, Kangran Zhao, Yize Chen 외

DeepFakes, which refer to AI-generated media content, have become an increasing concern due to their use as a means for disinformation. Detecting DeepFakes is currently solved with programmed machine learning algorithms.…

DeepFake DetectionExperimental DesignFace SwappingPrompt Engineering

Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes

2026-01-24 · Gautam Siddharth Kashyap, Harsh Joshi, Niharika Jain, Ebad Shabbir 외 arxiv

The rapid rise of deepfake technology poses a severe threat to social and political stability by enabling hyper-realistic synthetic media capable of manipulating public perception. However, existing detection methods str…

Contrastive LearningDeepFake Detection

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

2025-08-28 · Hao Tan, Jun Lan, Zichang Tan, Ajian Liu 외 arxiv

Deepfake detection remains a formidable challenge due to the complex and evolving nature of fake content in real-world scenarios. However, existing academic benchmarks suffer from severe discrepancies from industrial pra…

DeepFake Detection