paper-with-me

Papers

A Comparative Evaluation of Deep Learning Models for Speech Enhancement in Real-World Noisy Environments

2025-06-17 · Md Jahangir Alam Khondkar, Ajan Ahmed, Masudul Haider Imtiaz, Stephanie Schuckers

Speech enhancement, particularly denoising, is vital in improving the intelligibility and quality of speech signals for real-world applications, especially in noisy environments. While prior research has introduced various deep learning models for this purpose, many struggle to balance noise suppression, perceptual quality, and speaker-specific feature preservation, leaving a critical research gap in their comparative performance evaluation. This study benchmarks three state-of-the-art models Wave-U-Net, CMGAN, and U-Net, on diverse datasets such as SpEAR, VPQAD, and Clarkson datasets. These models were chosen due to their relevance in the literature and code accessibility. The evaluation reveals that U-Net achieves high noise suppression with SNR improvements of +71.96% on SpEAR, +64.83% on VPQAD, and +364.2% on the Clarkson dataset. CMGAN outperforms in perceptual quality, attaining the highest PESQ scores of 4.04 on SpEAR and 1.46 on VPQAD, making it well-suited for applications prioritizing natural and intelligible speech. Wave-U-Net balances these attributes with improvements in speaker-specific feature retention, evidenced by VeriSpeak score gains of +10.84% on SpEAR and +27.38% on VPQAD. This research indicates how advanced methods can optimize trade-offs between noise suppression, perceptual quality, and speaker recognition. The findings may contribute to advancing voice biometrics, forensic audio analysis, telecommunication, and speaker verification in challenging acoustic conditions.

📄 PDF Abstract BibTeX arXiv:2506.15000

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeaker RecognitionSpeaker VerificationSpeech Enhancement

Similar Papers 제목 키워드 기반

Lip-Reading Driven Deep Learning Approach for Speech Enhancement

2018-07-31 · Ahsan Adeel, Mandar Gogate, Amir Hussain, William M. Whitmer

This paper proposes a novel lip-reading driven deep learning framework for speech enhancement. The proposed approach leverages the complementary strengths of both deep learning and analytical acoustic modelling (filterin…

Acoustic ModellingDeep LearningLip Readingregression+1

Contextual Audio-Visual Switching For Speech Enhancement in Real-World Environments

2018-08-28 · Ahsan Adeel, Mandar Gogate, Amir Hussain

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only appr…

Lip ReadingSpeech Enhancement

Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness

2024-06-12 · Satyam Kumar, Sai Srujana Buddi, Utkarsh Oggy Sarawgi, Vineet Garg 외

Voice activity detection (VAD) is a critical component in various applications such as speech recognition, speech enhancement, and hands-free communication systems. With the increasing demand for personalized and context…

Action DetectionActivity DetectionSpeech Enhancementspeech-recognition+1

FNSE-SBGAN: Far-field Speech Enhancement with Schrodinger Bridge and Generative Adversarial Networks

2025-03-17 · Tong Lei, Qinwen Hu, Ziyao Lin, Andong Li 외

The prevailing method for neural speech enhancement predominantly utilizes fully-supervised deep learning with simulated pairs of far-field noisy-reverberant speech and clean speech. Nonetheless, these models frequently …

Speech Enhancement

Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems

2021-09-11 · Yangyang Xia, Buye Xu, Anurag Kumar

Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better r…

Speech EnhancementTriplet