paper-with-me

홈 › Papers

Speech Enhancement based on Denoising Autoencoder with Multi-branched Encoders

2020-01-06 · Cheng Yu, Ryandhimas E. Zezario, Jonathan Sherman, Yi-Yen Hsieh, Xugang Lu, Hsin-Min Wang, Yu Tsao

Deep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model generalizability to noisy conditions: (1) mismatched noisy condition during testing, i.e., the performance is generally sub-optimal when models are tested with unseen noise types that are not involved in the training data; (2) local focus on specific noisy conditions, i.e., models trained using multiple types of noises cannot optimally remove a specific noise type even though the noise type has been involved in the training data. These problems are common in real applications. In this paper, we propose a novel denoising autoencoder with a multi-branched encoder (termed DAEME) model to deal with these two problems. In the DAEME model, two stages are involved: offline and online. In the offline stage, we build multiple component models to form a multi-branched encoder based on a dynamically-sized decision tree(DSDT). The DSDT is built based on a prior knowledge of speech and noisy conditions (the speaker, environment, and signal factors are considered in this paper), where each component of the multi-branched encoder performs a particular mapping from noisy to clean speech along the branch in the DSDT. Finally, a decoder is trained on top of the multi-branched encoder. In the online stage, noisy speech is first processed by the tree and fed to each component model. The multiple outputs from these models are then integrated into the decoder to determine the final enhanced speech. Experimental results show that DAEME is superior to several baseline models in terms of objective evaluation metrics and the quality of subjective human listening tests.

📄 PDF Abstract BibTeX arXiv:2001.01538

Code (1)

WilliamYu1993/DAEME/tree/master/images

Tasks

DecoderDenoisingSpeech Enhancement

Methods 이 논문이 사용한 방법론

Denoising Autoencoder A Denoising Autoencoder is a modification on the autoencoder to prevent the network learning the identity function.…
Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

MIMO Speech Compression and Enhancement Based on Convolutional Denoising Autoencoder

2020-05-24 · You-Jin Li, Syu-Siang Wang, Yu Tsao, Borching Su

For speech-related applications in IoT environments, identifying effective methods to handle interference noises and compress the amount of data in transmissions is essential to achieve high-quality services. In this stu…

Denoising

Analysis of DNN Speech Signal Enhancement for Robust Speaker Recognition

2018-11-19

In this work, we present an analysis of a DNN-based autoencoder for speech enhancement, dereverberation and denoising. The target application is a robust speaker verification (SV) system. We start our approach by careful…

Data AugmentationDenoisingSpeaker RecognitionSpeaker Verification+1

Masked Autoencoders as Universal Speech Enhancer

2026-02-02 · Rajalaxmi Rajagopalan, Ritwik Giri, Zhiqiang Tang, Kyu Han arxiv

Supervised speech enhancement methods have been very successful. However, in practical scenarios, there is a lack of clean speech, and self-supervised learning-based (SSL) speech enhancement methods that offer comparable…

Self-Supervised LearningSpeech Enhancement

aTENNuate: Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio

2024-09-05 · Yan Ru Pei, Ritik Shrivastava, FNU Sidharth

We present aTENNuate, a simple deep state-space autoencoder configured for efficient online raw speech enhancement in an end-to-end fashion. The network's performance is primarily evaluated on raw speech denoising, with …

Audio DenoisingDenoisingSpeech DenoisingSpeech Enhancement+1

Improved far-field speech recognition using Joint Variational Autoencoder

2022-04-24 · Shashi Kumar, Shakti P. Rath, Abhishek Pandey

Automatic Speech Recognition (ASR) systems suffer considerably when source speech is corrupted with noise or room impulse responses (RIR). Typically, speech enhancement is applied in both mismatched and matched scenario …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Enhancement+2