paper-with-me

홈 › Papers

End-to-End Multi-Channel Speech Separation

2019-05-15 · Rongzhi Gu, Jian Wu, Shi-Xiong Zhang, Lian-Wu Chen, Yong Xu, Meng Yu, Dan Su, Yuexian Zou, Dong Yu

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech separation. The primary contributions of this work include 1) an integrated waveform-in waveform-out separation system in a single neural network architecture. 2) We reformulate the traditional short time Fourier transform (STFT) and inter-channel phase difference (IPD) as a function of time-domain convolution with a special kernel. 3) We further relaxed those fixed kernels to be learnable, so that the entire architecture becomes purely data-driven and can be trained from end-to-end. We demonstrate on the WSJ0 far-field speech separation task that, with the benefit of learnable spatial features, our proposed end-to-end multi-channel model significantly improved the performance of previous end-to-end single-channel method and traditional multi-channel methods.

📄 PDF Abstract BibTeX arXiv:1905.06286

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Efficient Integration of Multi-channel Information for Speaker-independent Speech Separation

2020-08-11

Although deep-learning-based methods have markedly improved the performance of speech separation over the past few years, it remains an open question how to integrate multi-channel signals for speech separation. We propo…

Deep ClusteringOpen-Ended Question AnsweringSpeech SeparationTransfer Learning

A comprehensive study of speech separation: spectrogram vs waveform separation

2019-05-17 · Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu, Shi-Xiong Zhang 외

Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently, a raw audio waveform separation network…

speech-recognitionSpeech RecognitionSpeech Separation

CasNet: Investigating Channel Robustness for Speech Separation

2022-10-27 · Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee, Yu Tsao 외

Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation performance, and cannot meet the requirement …

Speech Separation

Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments

2023-03-14 · Julian Neri, Sebastian Braun

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberat…

DecoderSpeech Separation

Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning

2020-03-09 · Rongzhi Gu, Shi-Xiong Zhang, Lian-Wu Chen, Yong Xu 외

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial fea…

Speech Separation