paper-with-me

홈 › Papers

MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models

2026-01-28 · Wenbo Xu, Wei Lu, Xiangyang Luo, Jiantao Zhou arxiv

Deepfake detection is a widely researched topic that is crucial for combating the spread of malicious content, with existing methods mainly modeling the problem as classification or spatial localization. The rapid advancements in generative models impose new demands on Deepfake detection. In this paper, we propose multimodal alignment and reinforcement for explainable Deepfake detection via vision-language models, termed MARE, which aims to enhance the accuracy and reliability of Vision-Language Models (VLMs) in Deepfake detection and reasoning. Specifically, MARE designs comprehensive reward functions, incorporating reinforcement learning from human feedback (RLHF), to incentivize the generation of text-spatially aligned reasoning content that adheres to human preferences. Besides, MARE introduces a forgery disentanglement module to capture intrinsic forgery traces from high-level facial semantics, thereby improving its authenticity detection capability. We conduct thorough evaluations on the reasoning content generated by MARE. Both quantitative and qualitative experimental results demonstrate that MARE achieves state-of-the-art performance in terms of accuracy and reliability.

📄 PDF Abstract BibTeX arXiv:2601.20433

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDeepFake Detection

Similar Papers 제목 키워드 기반

Explainable Deepfake Detection with RL Enhanced Self-Blended Images

2026-01-22 · Ning Jiang, Dingheng Zeng, Yanhong Liu, Haiyang Yi 외 arxiv

Most prior deepfake detection methods lack explainable outputs. With the growing interest in multimodal large language models (MLLMs), researchers have started exploring their use in interpretable deepfake detection. How…

Synthetic Data GenerationReinforcement LearningDomain GeneralizationDeepFake Detection

EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning

2025-10-18 · Haoran Sun, Chen Cai, Huiping Zhuang, Kong Aik Lee 외 arxiv

The rapid development of deepfake video technology has not only facilitated artistic creation but also made it easier to spread misinformation. Traditional deepfake video detection (DVD) methods face issues such as a lac…

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

2025-05-21 · Yuxuan Du, Zhendong Wang, Yuhao Luo, Caiyong Piao 외

The rapid emergence of multimodal deepfakes (visual and auditory content are manipulated in concert) undermines the reliability of existing detectors that rely solely on modality-specific artifacts or cross-modal inconsi…

cross-modal alignmentDeepFake DetectionFace Swapping

SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment

2026-03-26 · Sahibzada Adil Shahzad, Ammarah Hashmi, Junichi Yamagishi, Yusuke Yasuda 외 arxiv

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic…

Self-Supervised LearningDeepFake Detection

Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis

2025-10-23 · Lixiong Qin, Yang Zhang, Mei Wang, Jiani Hu 외 arxiv

The advancement of Multimodal Large Language Models (MLLMs) has bridged the gap between vision and language tasks, enabling the implementation of Explainable DeepFake Analysis (XDFA). However, current methods suffer from…

Multi-Task Learning