paper-with-me

Papers

Alethia: A Foundational Encoder for Voice Deepfakes

2026-04-30 · Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya Koppisetti arxiv

Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on $5$ different tasks with $56$ benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.

📄 PDF Abstract BibTeX arXiv:2605.00251

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationDeepFake Detection

Similar Papers 제목 키워드 기반

Vulnerability of Automatic Identity Recognition to Audio-Visual Deepfakes

2023-11-29 · Pavel Korshunov, Haolin Chen, Philip N. Garner, Sebastien Marcel

The task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. Ho…

Face RecognitionFace SwappingSpeaker Recognitiontext-to-speech+2

Combining Automatic Speaker Verification and Prosody Analysis for Synthetic Speech Detection

2022-10-31 · Luigi Attorresi, Davide Salvi, Clara Borrelli, Paolo Bestagini 외

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatic…

Audio CompressionFace SwappingRhythmSpeaker Verification+4

Joint Audio-Visual Deepfake Detection

2021-01-01 · ICCV 2021 10 · Yipin Zhou, Ser-Nam Lim

Deepfakes ("deep learning" + "fake") are synthetically-generated videos from AI algorithms. While they could be entertaining, they could also be misused for falsifying speeches and spreading misinformation. The proce…

DeepFake DetectionFace SwappingMisinformationtext-to-speech+2

Discussion Paper: The Threat of Real Time Deepfakes

2023-06-04 · Guy Frankovits, Yisroel Mirsky

Generative deep learning models are able to create realistic audio and video. This technology has been used to impersonate the faces and voices of individuals. These ``deepfakes'' are being used to spread misinformation,…

Misinformation

Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors

2021-09-07 · Hasam Khalid, Minha Kim, Shahroz Tariq, Simon S. Woo

Significant advancements made in the generation of deepfakes have caused security and privacy issues. Attackers can easily impersonate a person's identity in an image by replacing his face with the target person's face. …

DeepFake DetectionFace Swapping