Post-training for Deepfake Speech Detection
We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over 56,000 hours of genuine speech and 18,000 hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoints and source code are available online.
Code (1)
Tasks
Face SwappingSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection
Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervised pre-training objectives and spoof-specific artifacts. To a…
DeepFake DetectionData AugmentationSpoof DetectionFake Speech Wild: Detecting Deepfake Speech on Social Media Platform
The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on publ…
DeepFake DetectionData AugmentationPhoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause signifi…
DeepFake DetectionFace SwappingGraph Attentiontext-to-speech+1WavLM model ensemble for audio deepfake detection
Audio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious a…
Audio Deepfake DetectionData AugmentationDeepFake DetectionFace Swapping+3AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often un…
DeepFake DetectionSpeech Synthesis