paper-with-me

Papers

AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

2025-09-04 · Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie arxiv

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and highly diverse deepfake audio dataset comprising over 4,500 hours of synthetic audio generated by 11 recent TTS models and 10 vocoders, totalling 3 million clips. We further observe that most existing detectors default to binary supervised training, which can induce negative transfer across synthesis sources when the training data contains highly diverse deepfake patterns, impacting overall generalisation. As a complementary contribution, we propose an effective curriculum-learning-based approach to mitigate this effect. Extensive experiments show that existing detection models struggle to generalise to novel deepfakes and human speech in AUDETER, whereas XLR-based detectors trained on AUDETER achieve strong cross-domain performance across multiple benchmarks, achieving an EER of 1.87% on In-the-Wild. AUDETER is available on GitHub.

📄 PDF Abstract BibTeX arXiv:2509.04345

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionSpeech Synthesis

Similar Papers 제목 키워드 기반

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

2026-07-06 · Linxi Li, Yuncong Yu, Qianwei Guo, Liwei Jin 외 arxiv

While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects. Existing environmental audio datasets such as EnvSDD provide important initial r…

Audio Deepfake Detection

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

2025-05-25 · Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai 외

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental …

Audio Deepfake DetectionAudio GenerationBenchmarkingDeepFake Detection+1

Context and Transcripts Improve Detection of Deepfake Audios of Public Figures

2026-01-19 · Chongyang Gao, Marco Postiglione, Julian Baldwin, Natalia Denisenko 외 arxiv

Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provid…

The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio

2024-05-08 · Yuankun Xie, Yi Lu, Ruibo Fu, Zhengqi Wen 외

With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods. ALM-based deepfake audio currently exhibits widespread, high deception, and type versa…

Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping+2

AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset

2023-11-26 · Zhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat 외

The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detec…

DeepFake DetectionFace SwappingTemporal Forgery Localization