paper-with-me

Papers

Bts-e: Audio deepfake detection using breathing-talking-silence encoder

2023-05-05 · IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023 5 · Thien-Phuc Doan, Long Nguyen-Vu, Souhwan Jung, Kihun Hong

Voice phishing (vishing) is increasingly popular due to the development of speech synthesis technology. In particular, the use of deep learning to generate an arbitrary-content audio clip simulating the victim’s voice makes it difficult not only for humans but also for automatic speaker verification (ASV) systems to distinguish. Countermeasure (CM) systems have been developed recently to help ASV combat synthetic speech. In this work, we propose BTS-E, a framework to evaluate the correlation between Breathing, Talking (speech), and Silence sounds in an audio clip, then use this information for deepfake detection tasks. We argue that natural human sounds, such as breathing, are hard to synthesize by Text-to-speech (TTS) system. We conducted a large-scale evaluation using ASVspoof 2019 and 2021 evaluation set to validate our hypothesis. The experiment results show the applicability of the breathing sound feature in detecting deepfake voices. In general, the proposed system significantly increases the performance of the classifier by up to 46%.

📄 PDF Abstract BibTeX

Code (1)

josebeo2016/BTS-Encoder-ASVspoof pytorch

Tasks

Audio Deepfake DetectionDeepFake DetectionFace SwappingSpeaker VerificationSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning

2024-11-29 · CVPR 2025 1 · Stefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta Oneata

Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection - the focus of this paper. Here w…

BenchmarkingDeepFake DetectionFace Swapping

FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

2023-07-08 · Ganglai Wang, Peng Zhang, Junwen Xiong, Feihan Yang 외

DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake video detection is further improved. By on…

Face DetectionFace GenerationFace SwappingOptical Flow Estimation+1

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

2026-05-27 · Ke Liu, Jiwei Wei, Wenyu Zhang, Shuchang Zhou 외 arxiv

With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio-visual deepfake detection typically rely on cross-modal inconsistencies. In sing…

DeepFake Detection

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

2025-06-02 · CVPR 2025 1 · Yuan Gan, Jiaxu Miao, Yunze Wang, Yi Yang

Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of…

MisinformationTalking Head Generation

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection

2025-06-06 · Marcel Klemt, Carlotta Segna, Anna Rohrbach

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to…

BenchmarkingDeepFake DetectionFace SwappingMisinformation