paper-with-me

Papers

Semi-Supervised Multichannel Speech Enhancement With a Deep Speech Prior

2019-10-07 · IEEE/ACM Transactions on Audio, Speech, and Language Processing 2019 10 · Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii, Tatsuya Kawahara

This paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant called independent low-rank matrix analysis (ILRMA) have successfully been used for unsupervised speech enhancement, the low-rank assumption on the power spectral densities (PSDs) of all sources (speech and noise) does not hold in reality. To solve this problem, we replace a low-rank speech model with a deep generative speech model, i.e., formulate a probabilistic model of noisy speech by integrating a deep speech model, a low-rank noise model, and a full-rank or rank-1 model of spatial characteristics of speech and noise. The deep speech model is trained from clean speech data in an unsupervised auto-encoding variational Bayesian manner. Given multichannel noisy speech spectra, the full-rank or rank-1 spatial covariance matrices and PSDs of speech and noise are estimated in an unsupervised maximum-likelihood manner. Experimental results showed that the full-rank version of the proposed method was significantly better than MNMF, ILRMA, and the rank-1 version. We confirmed that the initialization-sensitivity and local-optimum problems of MNMF with many spatial parameters can be solved by incorporating the precise speech model.

📄 PDF Abstract BibTeX

Code (1)

sekiguchi92/SpeechEnhancement pytorch

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization

2018-11-16 · Simon Leglaive, Laurent Girin, Radu Horaud

In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neur…

Speech Enhancement

Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition

2019-03-22 · Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama 외

This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks

2020-10-21 · Yang Jiao

We propose a method for joint multichannel speech dereverberation with two spatial-aware tasks: direction-of-arrival (DOA) estimation and speech separation. The proposed method addresses involved tasks as a sequence to s…

Speech DereverberationSpeech EnhancementSpeech Separation

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

Speech enhancement using ego-noise references with a microphone array embedded in an unmanned aerial vehicle

2022-11-04 · Elisa Tengan, Thomas Dietzen, Santiago Ruiz, Mansour Alkmim 외

A method is proposed for performing speech enhancement using ego-noise references with a microphone array embedded in an unmanned aerial vehicle (UAV). The ego-noise reference signals are captured with microphones locate…

Speech Enhancement