Self-Supervised Video Forensics by Audio-Visual Anomaly Detection
Manipulated videos often contain subtle inconsistencies between their visual and audio signals. We propose a video forensics method, based on anomaly detection, that can identify these inconsistencies, and that can be trained solely using real, unlabeled data. We train an autoregressive model to generate sequences of audio-visual features, using feature sets that capture the temporal synchronization between video frames and sound. At test time, we then flag videos that the model assigns low probability. Despite being trained entirely on real videos, our model obtains strong performance on the task of detecting manipulated speech videos. Project site: https://cfeng16.github.io/audio-visual-forensics
Code (1)
Tasks
Anomaly DetectionDeepFake DetectionVideo ForensicsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leverag…
Representation LearningAV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Video Deepfake Detection
Multimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely d…
DeepFake DetectionFace SwappingSelf-Supervised LearningVideo ForensicsNPVForensics: Jointing Non-critical Phonemes and Visemes for Deepfake Detection
Deepfake technologies empowered by deep learning are rapidly evolving, creating new security concerns for society. Existing multimodal detection methods usually capture audio-visual inconsistencies to expose Deepfake vid…
DeepFake DetectionFace SwappingV2A-Mark: Versatile Deep Visual-Audio Watermarking for Manipulation Localization and Copyright Protection
AI-generated video has revolutionized short video production, filmmaking, and personalized media, making video local editing an essential tool. However, this progress also blurs the line between reality and fiction, posi…
Prompt LearningVideo EditingTelling Left from Right: Learning Spatial Correspondence of Sight and Sound
Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…
audio-visual learning