paper-with-me

홈 › Papers

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

2025-04-09 · Yuankun Xie, Ruibo Fu, Zhiyong Wang, Xiaopeng Wang, Songjun Cao, Long Ma, Haonan Cheng, Long Ye

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the alltype ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL frontend by learning specialized prompt tokens for ADD, requiring 458x fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types,we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets. The code is available online.

📄 PDF Abstract BibTeX arXiv:2504.06753

Code (1)

xieyuankun/all-type-add 공식 구현 pytorch

Tasks

AllAudio Deepfake DetectionAudio GenerationDeepFake DetectionFace SwappingSelf-Supervised Learning

Similar Papers 제목 키워드 기반

What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection

2023-12-15 · Xiaohui Zhang, Jiangyan Yi, Chenglong Wang, Chuyuan Zhang 외

The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Ex…

Audio Deepfake DetectionContinual LearningDeepFake DetectionFace Swapping+2

Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection

2026-01-02 · Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra, Abhijit Das 외 arxiv

While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored.…

Audio Deepfake Detection

WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection

2025-10-06 · Xi Xuan, Xuechen Liu, Wenxin Zhang, Yi-Cheng Lin 외 arxiv

Modern front-end design for speech deepfake detection relies on full fine-tuning of large pre-trained models like XLSR. However, this approach is not parameter-efficient and may lead to suboptimal generalization to reali…

DeepFake Detection

Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?

2024-08-20 · Yuankun Xie, Chenxu Xiong, Xiaopeng Wang, Zhiyong Wang 외

Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, genera…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

2026-04-09 · Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo 외 arxiv

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. …

Audio Deepfake Detection