Harder or Different? Understanding Generalization of Audio Deepfake Detection
Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS) models, i.e., are newer DeepFakes just 'harder' to detect? Or, is it because deepfakes generated with one model are fundamentally different to those generated using another model? We answer this question by decomposing the performance gap between in-domain and out-of-domain test data into 'hardness' and 'difference' components. Experiments performed using ASVspoof databases indicate that the hardness component is practically negligible, with the performance gap being attributed primarily to the difference component. This has direct implications for real-world deepfake detection, highlighting that merely increasing model capacity, the currently-dominant research trend, may not effectively address the generalization challenge.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio Deepfake DetectionDeepFake DetectionFace Swappingtext-to-speechText to SpeechMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Does Audio Deepfake Detection Generalize?
Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is o…
Audio Deepfake DetectionDeepFake DetectionFace Swappingtext-to-speech+1Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video
Recent advances in technology for hyper-realistic visual and audio effects provoke the concern that deepfake videos of political speeches will soon be indistinguishable from authentic video recordings. The conventional w…
Face SwappingHuman DetectionMisinformationtext-to-speech+1Audios Don't Lie: Multi-Frequency Channel Attention Mechanism for Audio Deepfake Detection
With the rapid development of artificial intelligence technology, the application of deepfake technology in the audio field has gradually increased, resulting in a wide range of security risks. Especially in the financia…
Audio Deepfake DetectionDeepFake DetectionFace SwappingAttack Agnostic Dataset: Towards Generalization and Stabilization of Audio DeepFake Detection
Audio DeepFakes allow the creation of high-quality, convincing utterances and therefore pose a threat due to its potential applications such as impersonation or fake news. Methods for detecting these manipulations should…
Audio Deepfake DetectionDeepFake DetectionFace SwappingXMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark
Recent advances in audio generation led to an increasing number of deepfakes, making the general public more vulnerable to financial scams, identity theft, and misinformation. Audio deepfake detectors promise to alleviat…
Audio GenerationFace SwappingMisinformation