paper-with-me

홈 › Papers

Single-Channel Speech Enhancement with Deep Complex U-Networks and Probabilistic Latent Space Models

2023-09-04 · Eike J. Nustede, Jörn Anemüller

In this paper, we propose to extend the deep, complex U-Network architecture for speech enhancement by incorporating a probabilistic (i.e., variational) latent space model. The proposed model is evaluated against several ablated versions of itself in order to study the effects of the variational latent space model, complex-value processing, and self-attention. Evaluation on the MS-DNS 2020 and Voicebank+Demand datasets yields consistently high performance. E.g., the proposed model achieves an SI-SDR of up to 20.2 dB, about 0.5 to 1.4 dB higher than its ablated version without probabilistic latent space, 2-2.4 dB higher than WaveUNet, and 6.7 dB above PHASEN. Compared to real-valued magnitude spectrogram processing with a variational U-Net, the complex U-Net achieves an improvement of up to 4.5 dB SI-SDR. Complex spectrum encoding as magnitude and phase yields best performance in anechoic conditions whereas real and imaginary part representation results in better generalization to (novel) reverberation conditions, possibly due to the underlying physics of sound.

📄 PDF Abstract BibTeX arXiv:2309.01535

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection

2022-06-20 · Yuan Chen, Yicheng Hsu, Mingsian R. Bai

Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressiv…

Action DetectionActivity DetectionSpeech Enhancement

Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders

2019-08-07 · Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin 외

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…

Speech Enhancement

Exploring the Potential of Data-Driven Spatial Audio Enhancement Using a Single-Channel Model

2024-04-22 · Arthur N. dos Santos, Bruno S. Masiero, Túlio C. L. Mateus

One key aspect differentiating data-driven single- and multi-channel speech enhancement and dereverberation methods is that both the problem formulation and complexity of the solutions are considerably more challenging i…

Direction of Arrival EstimationSpeech Enhancement

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

2022-03-15 · Zehua Zhang, Lu Zhang, Xuyi Zhuang, Yukun Qian 외

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time en…

DenoisingSpeech DenoisingSpeech Enhancement