paper-with-me

Papers

Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier

2023-12-13 · Yinlin Guo, Haofan Huang, Xi Chen, He Zhao, Yuehai Wang

With the rapid development of speech synthesis and voice conversion technologies, Audio Deepfake has become a serious threat to the Automatic Speaker Verification (ASV) system. Numerous countermeasures are proposed to detect this type of attack. In this paper, we report our efforts to combine the self-supervised WavLM model and Multi-Fusion Attentive classifier for audio deepfake detection. Our method exploits the WavLM model to extract features that are more conducive to spoofing detection for the first time. Then, we propose a novel Multi-Fusion Attentive (MFA) classifier based on the Attentive Statistics Pooling (ASP) layer. The MFA captures the complementary information of audio features at both time and layer levels. Experiments demonstrate that our methods achieve state-of-the-art results on the ASVspoof 2021 DF set and provide competitive results on the ASVspoof 2019 and 2021 LA set.

📄 PDF Abstract BibTeX arXiv:2312.08089

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake DetectionDeepFake Detection

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

WavLM model ensemble for audio deepfake detection

2024-08-14 · David Combei, Adriana Stan, Dan Oneata, Horia Cucu

Audio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious a…

Audio Deepfake DetectionData AugmentationDeepFake DetectionFace Swapping+3

Do Compact SSL Backbones Matter for Audio Deepfake Detection? A Controlled Study with RAPTOR

2026-03-06 · Ajinkya Kulkarni, Sandipana Dowerah, Atharva Kulkarni, Tanel Alumäe 외 arxiv

Self-supervised learning (SSL) underpins modern audio deepfake detection, yet most prior work centers on a single large wav2vec2-XLSR backbone, leaving compact under studied. We present RAPTOR, Representation Aware Pairw…

Self-Supervised LearningAudio Deepfake Detection

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection

2026-03-02 · Hashim Ali, Nithin Sai Adupa, Surya Subramani, Hafiz Malik arxiv

Self-supervised learning (SSL) has transformed speech processing, with benchmarks such as SUPERB establishing fair comparisons across diverse downstream tasks. Despite it's security-critical importance, Audio deepfake de…

Self-Supervised LearningAudio Deepfake Detection

Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System

2025-08-28 · Hashim Ali, Surya Subramani, Lekha Bollinani, Nithin Sai Adupa 외 arxiv

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-su…

Self-Supervised LearningAudio Deepfake Detection

Toward Noise-Aware Audio Deepfake Detection: Survey, SNR-Benchmarks, and Practical Recipes

2025-12-15 · Udayon Sen, Alka Luqman, Anupam Chattopadhyay arxiv

Deepfake audio detection has progressed rapidly with strong pre-trained encoders (e.g., WavLM, Wav2Vec2, MMS). However, performance in realistic capture conditions - background noise (domestic/office/transport), room rev…

Audio Deepfake Detection