paper-with-me

홈 › Papers

Retrieval-Augmented Audio Deepfake Detection

2024-04-22 · Zuheng Kang, Yayun He, Botao Zhao, Xiaoyang Qu, Junqing Peng, Jing Xiao, Jianzong Wang

With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse. However, most deepfake (DF) detection methods rely solely on the fuzzy knowledge learned by a single model, resulting in performance bottlenecks and transparency issues. Inspired by retrieval-augmented generation (RAG), we propose a retrieval-augmented detection (RAD) framework that augments test samples with similar retrieved samples for enhanced detection. We also extend the multi-fusion attentive classifier to integrate it with our proposed RAD framework. Extensive experiments show the superior performance of the proposed RAD framework over baseline methods, achieving state-of-the-art results on the ASVspoof 2021 DF set and competitive results on the 2019 and 2021 LA sets. Further sample analysis indicates that the retriever consistently retrieves samples mostly from the same speaker with acoustic characteristics highly consistent with the query audio, thereby improving detection performance.

📄 PDF Abstract BibTeX arXiv:2404.13892

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake DetectionDeepFake DetectionFace SwappingRAGRetrievalRetrieval-augmented GenerationSpeech Synthesistext-to-speechText to SpeechVoice Conversion

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Linguistically Augmented Audio Speech Data (LinguAS)

2026-06-08 · Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja arxiv

Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve. Yet, most detection models are trained to make inf…

Targeted Augmented Data for Audio Deepfake Detection

2024-07-10 · Marcella Astrid, Enjie Ghorbel, Djamila Aouada

The availability of highly convincing audio deepfake generators highlights the need for designing robust audio deepfake detectors. Existing works often rely solely on real and fake data available in the training set, whi…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

2026-05-19 · Aritra Marik, Marcel Klemt, Anna Rohrbach arxiv

With every advancement in generative AI models, forensics is under increasing pressure. The constant emergence of new generation techniques makes it impossible to collect data for each manipulation to train a deepfake de…

Emotion RecognitionDeepFake Detection

RAIDX: A Retrieval-Augmented Generation and GRPO Reinforcement Learning Framework for Explainable Deepfake Detection

2025-08-06 · Tianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He 외 arxiv

The rapid advancement of AI-generation models has enabled the creation of hyperrealistic imagery, posing ethical risks through widespread misinformation. Current deepfake detection methods, categorized as face specific d…

Reinforcement LearningDeepFake Detection

Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?

2024-08-20 · Yuankun Xie, Chenxu Xiong, Xiaopeng Wang, Zhiyong Wang 외

Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, genera…

Audio Deepfake DetectionDeepFake DetectionFace Swapping