paper-with-me

홈 › Papers

Watch Those Words: Video Falsification Detection Using Word-Conditioned Facial Motion

2021-12-21 · Shruti Agarwal, Liwen Hu, Evonne Ng, Trevor Darrell, Hao Li, Anna Rohrbach

In today's era of digital misinformation, we are increasingly faced with new threats posed by video falsification techniques. Such falsifications range from cheapfakes (e.g., lookalikes or audio dubbing) to deepfakes (e.g., sophisticated AI media synthesis methods), which are becoming perceptually indistinguishable from real videos. To tackle this challenge, we propose a multi-modal semantic forensic approach to discover clues that go beyond detecting discrepancies in visual quality, thereby handling both simpler cheapfakes and visually persuasive deepfakes. In this work, our goal is to verify that the purported person seen in the video is indeed themselves by detecting anomalous facial movements corresponding to the spoken words. We leverage the idea of attribution to learn person-specific biometric patterns that distinguish a given speaker from others. We use interpretable Action Units (AUs) to capture a person's face and head movement as opposed to deep CNN features, and we are the first to use word-conditioned facial motion analysis. We further demonstrate our method's effectiveness on a range of fakes not seen in training including those without video manipulation, that were not addressed in prior work.

📄 PDF Abstract BibTeX arXiv:2112.10936

Code (1)

agarwalshruti15/wtw_project_page 공식 구현

Tasks

Misinformation

Similar Papers 제목 키워드 기반

The Compositional Nature of Event Representations in the Human Brain

2015-05-25

How does the human brain represent simple compositions of constituents: actors, verbs, objects, directions, and locations? Subjects viewed videos during neuroimaging (fMRI) sessions from which sentential descriptions of …

Classification

A Cooperative Statistical Approach for Abnormal Node Detection with Adversary Resistance

2023-11-28 · Yingying Huangfu, Tian Bai

Distinguishing abnormal nodes from those with normal packet loss in clusters helps reduce the loss of clustered network resources. The detection performance of existing detection schemes is limited by the techniques to q…

Counteracting Duration Bias in Video Recommendation via Counterfactual Watch Time

2024-06-12 · Haiyuan Zhao, Guohao Cai, Jieming Zhu, Zhenhua Dong 외

In video recommendation, an ongoing effort is to satisfy users' personalized information needs by leveraging their logged watch time. However, watch time prediction suffers from duration bias, hindering its ability to re…

counterfactualRecommendation Systems

Active Light Modulation to Counter Manipulation of Speech Visual Content

2025-04-30 · Hadleigh Schwartz, Xiaofeng Yan, Charles J. Carver, Xia Zhou

High-profile speech videos are prime targets for falsification, owing to their accessibility and influence. This work proposes Spotlight, a low-overhead and unobtrusive system for protecting live speech videos from visua…

Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection

2025-01-05 · Sung Jin Um, DongJin Kim, Sangmin Lee, Jung Uk Kim

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent w…

Contrastive LearningHighlight DetectionMoment RetrievalRetrieval