paper-with-me

Papers

Evaluating the Impact of Discriminative and Generative E2E Speech Enhancement Models on Syllable Stress Preservation

2024-12-11 · Rangavajjala Sankara Bharadwaj, Jhansi Mallela, Sai Harshitha Aluru, Chiranjeevi Yarra

Automatic syllable stress detection is a crucial component in Computer-Assisted Language Learning (CALL) systems for language learners. Current stress detection models are typically trained on clean speech, which may not be robust in real-world scenarios where background noise is prevalent. To address this, speech enhancement (SE) models, designed to enhance speech by removing noise, might be employed, but their impact on preserving syllable stress patterns is not well studied. This study examines how different SE models, representing discriminative and generative modeling approaches, affect syllable stress detection under noisy conditions. We assess these models by applying them to speech data with varying signal-to-noise ratios (SNRs) from 0 to 20 dB, and evaluating their effectiveness in maintaining stress patterns. Additionally, we explore different feature sets to determine which ones are most effective for capturing stress patterns amidst noise. To further understand the impact of SE models, a human-based perceptual study is conducted to compare the perceived stress patterns in SE-enhanced speech with those in clean speech, providing insights into how well these models preserve syllable stress as perceived by listeners. Experiments are performed on English speech data from non-native speakers of German and Italian. And the results reveal that the stress detection performance is robust with the generative SE models when heuristic features are used. Also, the observations from the perceptual study are consistent with the stress detection outcomes under all SE models.

📄 PDF Abstract BibTeX arXiv:2412.08306

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Analysing Diffusion-based Generative Approaches versus Discriminative Approaches for Speech Restoration

2022-11-04 · Jean-Marie Lemercier, Julius Richter, Simon Welker, Timo Gerkmann

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks l…

Bandwidth ExtensionSpeech DenoisingSpeech DereverberationSpeech Enhancement

GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning

2024-10-17 · Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain gen…

Generative Adversarial NetworkSpeech Enhancement

Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement

2024-09-15 · Yudong Yang, Zhan Liu, Wenyi Yu, Guangzhi Sun 외

Diffusion-based generative models have recently achieved remarkable results in speech and vocal enhancement due to their ability to model complex speech data distributions. While these models generalize well to unseen ac…

Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

2024-12-19 · Leying Zhang, Wangyou Zhang, Chenda Li, Yanmin Qian

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been und…

AttributeSpeech Enhancementtext-to-speechText to Speech

Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement

2024-06-19 · Chenda Li, Samuele Cornell, Shinji Watanabe, Yanmin Qian

Diffusion-based generative models (DGMs) have recently attracted attention in speech enhancement research (SE) as previous works showed a remarkable generalization capability. However, DGMs are also computationally inten…

Speech Enhancement