paper-with-me

홈 › Papers

Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments

2022-07-15 · Kouhei Sekiguchi, Aditya Arie Nugraha, Yicheng Du, Yoshiaki Bando, Mathieu Fontaine, Kazuyoshi Yoshii

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source separation method called fast multichannel nonnegative matrix factorization (FastMNMF) that works well in various environments thanks to its unsupervised nature. Its heavy computational cost, however, prevents its application to real-time processing. In contrast, a supervised beamforming method that uses a deep neural network (DNN) for estimating spatial information of speech and noise readily fits real-time processing, but suffers from drastic performance degradation in mismatched conditions. Given such complementary characteristics, we propose a dual-process robust online speech enhancement method based on DNN-based beamforming with FastMNMF-guided adaptation. FastMNMF (back end) is performed in a mini-batch style and the noisy and enhanced speech pairs are used together with the original parallel training data for updating the direction-aware DNN (front end) with backpropagation at a computationally-allowable interval. This method is used with a blind dereverberation method called weighted prediction error (WPE) for transcribing the noisy reverberant speech of a speaker, which can be detected from video or selected by a user's hand gesture or eye gaze, in a streaming manner and spatially showing the transcriptions with an AR technique. Our experiment showed that the word error rate was improved by more than 10 points with the run-time adaptation using only twelve minutes of observation.

📄 PDF Abstract BibTeX arXiv:2207.07296

Code (1)

sekiguchi92/SpeechEnhancement pytorch

Tasks

blind source separationSpeech Enhancement

Similar Papers 제목 키워드 기반

Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments

2022-07-15 · Yicheng Du, Aditya Arie Nugraha, Kouhei Sekiguchi, Yoshiaki Bando 외

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simula…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Distant Speech RecognitionNoisy Speech Recognition+3

MetricGAN-OKD: Multi-Metric Optimization of MetricGAN via Online Knowledge Distillation for Speech Enhancement

2023-07-24 · ICML 2023 7 · WooSeok Shin, Byung Hoon Lee, Jin Sob Kim, Hyun Joon Park 외

In speech enhancement, MetricGAN-based approaches reduce the discrepancy between the $L_p$ loss and evaluation metrics by utilizing a non-differentiable evaluation metric as the objective function. However, optimizing mu…

Knowledge DistillationSpeech Enhancement

BERT for Joint Multichannel Speech Dereverberation with Spatial-aware Tasks

2020-10-21 · Yang Jiao

We propose a method for joint multichannel speech dereverberation with two spatial-aware tasks: direction-of-arrival (DOA) estimation and speech separation. The proposed method addresses involved tasks as a sequence to s…

Speech DereverberationSpeech EnhancementSpeech Separation

Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement

2021-08-27 · Yuzi Yan, Wei-Qiang Zhang, Michael T. Johnson

As the cornerstone of other important technologies, such as speech recognition and speech synthesis, speech enhancement is a critical area in audio signal processing. In this paper, a new deep learning structure for spee…

Audio Signal ProcessingSpeech Enhancementspeech-recognitionSpeech Recognition+1

ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding

2022-07-19 · Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang 외

This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancement+4