paper-with-me

Papers

Sequential Multi-Frame Neural Beamforming for Speech Separation and Enhancement

2019-11-18 · Zhong-Qiu Wang, Hakan Erdogan, Scott Wisdom, Kevin Wilson, Desh Raj, Shinji Watanabe, Zhuo Chen, John R. Hershey

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture trained with a novel stabilized signal-to-noise ratio loss function. For beamforming, we explore multiple ways of computing time-varying covariance matrices, including factorizing the spatial covariance into a time-varying amplitude component and a time-invariant spatial component, as well as using block-based techniques. In addition, we introduce a multi-frame beamforming method which improves the results significantly by adding contextual frames to the beamforming formulations. We extensively evaluate and analyze the effects of window size, block size, and multi-frame context size for these methods. Our best method utilizes a sequence of three neural separation and multi-frame time-invariant spatial beamforming stages, and demonstrates an average improvement of 2.75 dB in scale-invariant signal-to-noise ratio and 14.2% absolute reduction in a comparative speech recognition metric across four challenging reverberant speech enhancement and separation tasks. We also use our three-speaker separation model to separate real recordings in the LibriCSS evaluation set into non-overlapping tracks, and achieve a better word error rate as compared to a baseline mask based beamformer.

📄 PDF Abstract BibTeX arXiv:1911.07953

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker SeparationSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel Output

2021-02-05 · Hangting Chen, Yang Yi, Dang Feng, Pengyuan Zhang

Time-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS). Classic multi-channel speech processing framework employs signal estimation and beamforming. For example…

blind source separationSpeech Separation

Multi-microphone Complex Spectral Mapping for Utterance-wise and Continuous Speech Separation

2020-10-04 · Zhong-Qiu Wang, Peidong Wang, DeLiang Wang

We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation an…

Speaker SeparationSpeech Separation

MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation

2022-12-07 · Yanjie Fu, Haoran Yin, Meng Ge, Longbiao Wang 외

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speaker feature, face image or directional in…

Speech Separation

Exploring the Integration of Speech Separation and Recognition with Self-Supervised Learning Representation

2023-07-23 · Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang, Samuele Cornell 외

Neural speech separation has made remarkable progress and its integration with automatic speech recognition (ASR) is an important direction towards realizing multi-speaker ASR. This work provides an insightful investigat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Self-Supervised LearningSpeaker Recognition+3

DNN-Free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online FastMNMF

2022-07-22 · Aditya Arie Nugraha, Kouhei Sekiguchi, Mathieu Fontaine, Yoshiaki Bando 외

This paper describes a practical dual-process speech enhancement system that adapts environment-sensitive frame-online beamforming (front-end) with help from environment-free block-online source separation (back-end). To…

blind source separationSpeech Enhancement