paper-with-me

Papers

VoiceWukong: Benchmarking Deepfake Voice Detection

2024-09-10 · Ziwei Yan, Yanjie Zhao, Haoyu Wang

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for evaluating detectors. Existing datasets are limited in language diversity and lack many manipulations encountered in real-world production environments. To fill this gap, we propose VoiceWukong, a benchmark designed to evaluate the performance of deepfake voice detectors. To build the dataset, we first collected deepfake voices generated by 19 advanced and widely recognized commercial tools and 15 open-source tools. We then created 38 data variants covering six types of manipulations, constructing the evaluation dataset for deepfake voice detection. VoiceWukong thus includes 265,200 English and 148,200 Chinese deepfake voice samples. Using VoiceWukong, we evaluated 12 state-of-the-art detectors. AASIST2 achieved the best equal error rate (EER) of 13.50%, while all others exceeded 20%. Our findings reveal that these detectors face significant challenges in real-world applications, with dramatically declining performance. In addition, we conducted a user study with more than 300 participants. The results are compared with the performance of the 12 detectors and a multimodel large language model (MLLM), i.e., Qwen2-Audio, where different detectors and humans exhibit varying identification capabilities for deepfake voices at different deception levels, while the LALM demonstrates no detection ability at all. Furthermore, we provide a leaderboard for deepfake voice detection, publicly available at {https://voicewukong.github.io}.

📄 PDF Abstract BibTeX arXiv:2409.06348

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingFace SwappingLarge Language Modeltext-to-speechText to SpeechVoice Conversion

Similar Papers 제목 키워드 기반

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

2025-05-25 · Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai 외

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental …

Audio Deepfake DetectionAudio GenerationBenchmarkingDeepFake Detection+1

AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations

2025-07-28 · Zhixi Cai, Kartik Kuckreja, Shreya Ghosh, Akanksha Chuchra 외 arxiv

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturba…

Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024

2024-09-03 · Anmol Guragain, Tianchi Liu, Zihan Pan, Hardik B. Sailor 외

This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generat…

DeepFake DetectionFace SwappingVoice Anti-spoofing

SingFake: Singing Voice Deepfake Detection

2023-09-14 · Yongyi Zang, You Zhang, Mojtaba Heydari, Zhiyao Duan

The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs c…

DeepFake DetectionFace SwappingSinging Voice SynthesisSynthetic Speech Detection

FSD: An Initial Chinese Dataset for Fake Song Detection

2023-09-05 · Yuankun Xie, Jingjing Zhou, Xiaolin Lu, Zhenghao Jiang 외

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authentic…

Audio Deepfake DetectionDeepFake DetectionFace SwappingFake Song Detection+2