paper-with-me

Papers

Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments

2026-03-16 · Anacin, Angela, Shruti Kshirsagar, Anderson R. Avila arxiv

Logical Access (LA) attacks, also known as audio deepfake attacks, use Text-to-Speech (TTS) or Voice Conversion (VC) methods to generate spoofed speech data. This can represent a serious threat to Automatic Speaker Verification (ASV) systems, as intruders can use such attacks to bypass voice biometric security. In this study, we investigate the correlation between speech quality and the performance of audio spoofing detection systems (i.e., LA task). For that, the performance of two enhancement algorithms is evaluated based on two perceptual speech quality measures, namely Perceptual Evaluation of Speech Quality (PESQ) and Speech-to-Reverberation Modulation Ratio (SRMR), and in respect to their impact on the audio spoofing detection system. We adopted the LA dataset, provided in the ASVspoof 2019 Challenge, and corrupted its test set with different Signal-to-Noise Ratio (SNR) levels, while leaving the training data untouched. Enhancement was applied to attenuate the detrimental effects of noisy speech, and the performances of two models, Speech Enhancement Generative Adversarial Network (SEGAN) and Metric-Optimized Generative Adversarial Network Plus (MetricGAN+), were compared. Although we expect that speech quality will correlate well with speech applications' performance, it can also have as a side effect on downstream tasks if unwanted artifacts are introduced or relevant information is removed from the speech signal. Our results corroborate with this hypothesis, as we found that the enhancement algorithm leading to the highest speech quality scores, MetricGAN+, provided the lowest Equal Error Rate (EER) on the audio spoofing detection task, whereas the enhancement method with the lowest speech quality scores, SEGAN, led to the lowest EER, thus leading to better performance on the LA task.

📄 PDF Abstract BibTeX arXiv:2603.14767

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Deepfake DetectionSpeaker VerificationSpeech EnhancementVoice Conversion

Similar Papers 제목 키워드 기반

AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds

2025-09-04 · Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani 외 arxiv

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often un…

DeepFake DetectionSpeech Synthesis

Realism to Deception: Investigating Deepfake Detectors Against Face Enhancement

2025-09-08 · Muhammad Saad Saeed, Ijaz Ul Haq, Khalid Malik arxiv

Face enhancement techniques are widely used to enhance facial appearance. However, they can inadvertently distort biometric features, leading to significant decrease in the accuracy of deepfake detectors. This study hypo…

Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake

2024-03-31 · Orchid Chetia Phukan, Gautam Siddharth Kashyap, Arun Balaji Buduru, Rajesh Sharma

In this work, we investigate multilingual speech Pre-Trained models (PTMs) for Audio deepfake detection (ADD). We hypothesize that multilingual PTMs trained on large-scale diverse multilingual data gain knowledge about d…

Audio Deepfake DetectionDeepFake DetectionEmotion RecognitionFace Swapping

Investigating Prosodic Signatures via Speech Pre-Trained Models for Audio Deepfake Source Attribution

2024-12-23 · Orchid Chetia Phukan, Drishti Singh, Swarup Ranjan Behera, Arun Balaji Buduru 외

In this work, we investigate various state-of-the-art (SOTA) speech pre-trained models (PTMs) for their capability to capture prosodic signatures of the generative sources for audio deepfake source attribution (ADSD). Th…

Audio Deepfake DetectionDeepFake DetectionFace SwappingSpeaker Recognition+2

Investigating self-supervised representations for audio-visual deepfake detection

2025-11-21 · Dragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta Oneata arxiv

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried with…

DeepFake Detection