paper-with-me

홈 › Papers

Consistency-aware multi-channel speech enhancement using deep neural networks

2020-02-14

This paper proposes a deep neural network (DNN)-based multi-channel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filtering can be efficiently implemented in the T-F domain. In such a case, ordinary objective functions are computed on the estimated T-F mask or spectrogram. However, the estimated spectrogram is often inconsistent, and its amplitude and phase may change when the spectrogram is converted back to the time-domain. That is, the objective function does not evaluate the enhanced time-domain signal properly. To address this problem, we propose to use an objective function defined on the reconstructed time-domain signal. Specifically, speech enhancement is conducted by multi-channel Wiener filtering in the T-F domain, and its result is converted back to the time-domain. We propose two objective functions computed on the reconstructed signal where the first one is defined in the time-domain, and the other one is defined in the T-F domain. Our experiment demonstrates the effectiveness of the proposed system comparing to T-F masking and mask-based beamforming.

📄 PDF Abstract BibTeX arXiv:2002.05831

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks

2020-10-21 · Yang Jiao

We propose a method for joint multichannel speech dereverberation with two spatial-aware tasks: direction-of-arrival (DOA) estimation and speech separation. The proposed method addresses involved tasks as a sequence to s…

Speech DereverberationSpeech EnhancementSpeech Separation

A Training Framework for Stereo-Aware Speech Enhancement using Deep Neural Networks

2021-12-09 · Bahareh Tolooshams, Kazuhito Koishida

Deep learning-based speech enhancement has shown unprecedented performance in recent years. The most popular mono speech enhancement frameworks are end-to-end networks mapping the noisy mixture into an estimate of the cl…

Deep LearningSpeech Enhancement

Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions

2021-10-27 · Wangyou Zhang, Jing Shi, Chenda Li, Shinji Watanabe 외

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on the time-domain speech enhancement model …

Speech Enhancementspeech-recognitionSpeech Recognition

A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement

2021-08-01 · INTERSPEECH 2021 2021 8 · Xinlei Ren, Xu Zhang, LianWu Chen, Xiguang Zheng 외

People are meeting through video conferencing more often. While single channel speech enhancement techniques are useful for the individual participants, the speech quality will be significantly degraded in large meeting …

CPUSpeech Enhancement