paper-with-me

Papers

AvatarShield: Visual Reinforcement Learning for Human-Centric Video Forgery Detection

2025-05-21 · Zhipei Xu, Xuanyu Zhang, Xing Zhou, Jian Zhang

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, particularly in video generation, has led to unprecedented creative capabilities but also increased threats to information integrity, identity security, and public trust. Existing detection methods, while effective in general scenarios, lack robust solutions for human-centric videos, which pose greater risks due to their realism and potential for legal and ethical misuse. Moreover, current detection approaches often suffer from poor generalization, limited scalability, and reliance on labor-intensive supervised fine-tuning. To address these challenges, we propose AvatarShield, the first interpretable MLLM-based framework for detecting human-centric fake videos, enhanced via Group Relative Policy Optimization (GRPO). Through our carefully designed accuracy detection reward and temporal compensation reward, it effectively avoids the use of high-cost text annotation data, enabling precise temporal modeling and forgery detection. Meanwhile, we design a dual-encoder architecture, combining high-level semantic reasoning and low-level artifact amplification to guide MLLMs in effective forgery detection. We further collect FakeHumanVid, a large-scale human-centric video benchmark that includes synthesis methods guided by pose, audio, and text inputs, enabling rigorous evaluation of detection methods in real-world scenes. Extensive experiments show that AvatarShield significantly outperforms existing approaches in both in-domain and cross-domain detection, setting a new standard for human-centric video forensics.

📄 PDF Abstract BibTeX arXiv:2505.15173

Code (1)

zhipeixu/fakeshield pytorch

Tasks

reinforcement-learningReinforcement Learningtext annotationVideo ForensicsVideo Generation

Similar Papers 제목 키워드 기반

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

2025-11-25 · Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu 외 arxiv

While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This high-level v…

Reinforcement Learning

EgoVLM: Policy Optimization for Egocentric Video Understanding

2025-06-03 · Ashwin Vinod, Shrey Pandit, Aditya Vavre, Linshen Liu

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically…

EgoSchemaQuestion Answeringreinforcement-learningReinforcement Learning+2

Use of Affective Visual Information for Summarization of Human-Centric Videos

2021-07-08 · Berkay Köprü, Engin Erzin

Increasing volume of user-generated human-centric video content and their applications, such as video retrieval and browsing, require compact representations that are addressed by the video summarization literature. Curr…

Emotion RecognitionRetrievalSupervised Video SummarizationVideo Retrieval+1

Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues

2025-01-01 · CVPR 2025 1 · Sihong Huang, Jiaxin Wu, XiaoYong Wei, Yi Cai 외

Understanding human behavior and the environmental information in the egocentric video is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-perso…

Action RecognitionScene RecognitionVideo Alignment

Ego-Pose Estimation and Forecasting as Real-Time PD Control

2019-06-07 · ICCV 2019 10 · Ye Yuan, Kris Kitani

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented…

Egocentric Pose EstimationHuman Pose ForecastingPose EstimationReinforcement Learning+2