paper-with-me

홈 › Papers

Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition

2025-11-27 · Maheswar Bora, Tashvik Dhamija, Shukesh Reddy, Baptiste Chopin, Pranav Balaji, Abhijit Das, Antitza Dantcheva arxiv

Deepfake generation has witnessed remarkable progress, contributing to highly realistic generated images, videos, and audio. While technically intriguing, such progress has raised serious concerns related to the misuse of manipulated media. To mitigate such misuse, robust and reliable deepfake detection is urgently needed. Towards this, we propose a novel network FauxNet, which is based on pre-trained Visual Speech Recognition (VSR) features. By extracting temporal VSR features from videos, we identify and segregate real videos from manipulated ones. The holy grail in this context has to do with zero-shot detection, i.e., generalizable detection, which we focus on in this work. FauxNet consistently outperforms the state-of-the-art in this setting. In addition, FauxNet is able to attribute - distinguish between generation techniques from which the videos stem. Finally, we propose new datasets, referred to as Authentica-Vox and Authentica-HDTF, comprising about 38,000 real and fake videos in total, the latter created with six recent deepfake generation techniques. We provide extensive analysis and results on the Authentica datasets and FaceForensics++, demonstrating the superiority of FauxNet. The Authentica datasets will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2511.22443

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Speech RecognitionDeepFake Detection

Similar Papers 제목 키워드 기반

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

2026-08-24 · Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke 외 arxiv

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigate…

DeepFake DetectionFace Parsing

Generalizable speech deepfake detection via meta-learned LoRA

2025-02-15 · Janne Laakkonen, Ivan Kukanov, Ville Hautamäki

Generalizable deepfake detection can be formulated as a detection problem where labels (bonafide and fake) are fixed but distributional drift affects the deepfake set. We can always train our detector with one-selected a…

DeepFake DetectionFace SwappingMeta-Learning

Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video

2022-02-25 · Matthew Groh, Aruna Sankaranarayanan, Nikhil Singh, Dong Young Kim 외

Recent advances in technology for hyper-realistic visual and audio effects provoke the concern that deepfake videos of political speeches will soon be indistinguishable from authentic video recordings. The conventional w…

Face SwappingHuman DetectionMisinformationtext-to-speech+1

Joint Audio-Visual Deepfake Detection

2021-01-01 · ICCV 2021 10 · Yipin Zhou, Ser-Nam Lim

Deepfakes ("deep learning" + "fake") are synthetically-generated videos from AI algorithms. While they could be entertaining, they could also be misused for falsifying speeches and spreading misinformation. The proce…

DeepFake DetectionFace SwappingMisinformationtext-to-speech+2

Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection

2025-09-17 · Janne Laakkonen, Ivan Kukanov, Ville Hautamäki arxiv

Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail…

Audio Deepfake DetectionRepresentation Learning