paper-with-me

Papers

All-neural beamformer for continuous speech separation

2021-10-13 · Zhuohuang Zhang, Takuya Yoshioka, Naoyuki Kanda, Zhuo Chen, Xiaofei Wang, Dongmei Wang, Sefik Emre Eskimez

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is transcribing a meeting conversation recorded by a microphone array. Prior studies explored various deep learning models for time-frequency mask estimation, followed by a minimum variance distortionless response (MVDR) filter to improve the automatic speech recognition (ASR) accuracy. The performance of these methods is fundamentally upper-bounded by MVDR's spatial selectivity. Recently, the all deep learning MVDR (ADL-MVDR) model was proposed for neural beamforming and demonstrated superior performance in a target speech extraction task using pre-segmented input. In this paper, we further adapt ADL-MVDR to the CSS task with several enhancements to enable end-to-end neural beamforming. The proposed system achieves significant word error rate reduction over a baseline spectral masking system on the LibriCSS dataset. Moreover, the proposed neural beamformer is shown to be comparable to a state-of-the-art MVDR-based system in real meeting transcription tasks, including AMI, while showing potentials to further simplify the runtime implementation and reduce the system latency with frame-wise processing.

📄 PDF Abstract BibTeX arXiv:2110.06428

Code (0)

등록된 구현이 없습니다.

Tasks

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Extractionspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Low-Latency Speaker-Independent Continuous Speech Separation

2019-04-13 · Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao 외

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of whi…

speech-recognitionSpeech RecognitionSpeech Separation

Enhanced Neural Beamformer with Spatial Information for Target Speech Extraction

2023-06-28 · Aoqi Guo, Junnan Wu, Peng Gao, Wenbo Zhu 외

Recently, deep learning-based beamforming algorithms have shown promising performance in target speech extraction tasks. However, most systems do not fully utilize spatial information. In this paper, we propose a target …

Dimensionality ReductionSpeech ExtractionSpeech Separation

Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation

2023-05-18 · Yanjie Fu, Meng Ge, Honglong Wang, Nan Li 외

Recently, stunning improvements on multi-channel speech separation have been achieved by neural beamformers when direction information is available. However, most of them neglect to utilize speaker's 2-dimensional (2D) l…

AllSpeech Separation

WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation

2020-11-18

This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation

2021-04-17 · Xiyun Li, Yong Xu, Meng Yu, Shi-Xiong Zhang 외

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1