Bts-e: Audio deepfake detection using breathing-talking-silence encoder
Voice phishing (vishing) is increasingly popular due to the development of speech synthesis technology. In particular, the use of deep learning to generate an arbitrary-content audio clip simulating the victim’s voice makes it difficult not only for humans but also for automatic speaker verification (ASV) systems to distinguish. Countermeasure (CM) systems have been developed recently to help ASV combat synthetic speech. In this work, we propose BTS-E, a framework to evaluate the correlation between Breathing, Talking (speech), and Silence sounds in an audio clip, then use this information for deepfake detection tasks. We argue that natural human sounds, such as breathing, are hard to synthesize by Text-to-speech (TTS) system. We conducted a large-scale evaluation using ASVspoof 2019 and 2021 evaluation set to validate our hypothesis. The experiment results show the applicability of the breathing sound feature in detecting deepfake voices. In general, the proposed system significantly increases the performance of the classifier by up to 46%.
Code (1)
Tasks
Audio Deepfake DetectionDeepFake DetectionFace SwappingSpeaker VerificationSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection - the focus of this paper. Here w…
BenchmarkingDeepFake DetectionFace SwappingFTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction
DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By on…
Face DetectionFace GenerationFace SwappingOptical Flow Estimation+1From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In sing…
DeepFake DetectionSilence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of…
MisinformationTalking Head GenerationDeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to…
BenchmarkingDeepFake DetectionFace SwappingMisinformation