Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
Currently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs. These ALMs have significantly lowered the barrier to creating deepfake audio, generating highly realistic and diverse types of deepfake audio, which pose severe threats to society. Consequently, effective audio deepfake detection technologies to detect ALM-based audio have become increasingly critical. This paper investigate the effectiveness of current countermeasure (CM) against ALM-based audio. Specifically, we collect 12 types of the latest ALM-based deepfake audio and utilizing the latest CMs to evaluate. Our findings reveal that the latest codec-trained CM can effectively detect ALM-based audio, achieving 0% equal error rate under most ALM test conditions, which exceeded our expectations. This indicates promising directions for future research in ALM-based deepfake audio detection.
Code (1)
Tasks
Audio Deepfake DetectionDeepFake DetectionFace SwappingSimilar Papers 제목 키워드 기반
The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods. ALM-based deepfake audio currently exhibits widespread, high deception, and type versa…
Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping+2Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
With the proliferation of deepfake audio, there is an urgent need to investigate their attribution. Current source tracing methods can effectively distinguish in-distribution (ID) categories. However, the rapid evolution…
Audio Deepfake DetectionDeepFake DetectionFace SwappingCodecfake: An Initial Dataset for Detecting LLM-based Deepfake Audio
With the proliferation of Large Language Model (LLM) based deepfake audio, there is an urgent need for effective detection methods. Previous deepfake audio generation methods typically involve a multi-step generation pro…
Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping+3Does Audio Deepfake Detection Generalize?
Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is o…
Audio Deepfake DetectionDeepFake DetectionFace Swappingtext-to-speech+1Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier
Generative AI technologies, including text-to-speech (TTS) and voice conversion (VC), frequently become indistinguishable from genuine samples, posing challenges for individuals in discerning between real and syntheti…
Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping+3