paper-with-me

홈 › Papers

Enhancing End-to-End Multi-channel Speech Separation via Spatial Feature Learning

2020-03-09 · Rongzhi Gu, Shi-Xiong Zhang, Lian-Wu Chen, Yong Xu, Meng Yu, Dan Su, Yuexian Zou, Dong Yu

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In this work, we propose an integrated architecture for learning spatial features directly from the multi-channel speech waveforms within an end-to-end speech separation framework. In this architecture, time-domain filters spanning signal channels are trained to perform adaptive spatial filtering. These filters are implemented by a 2d convolution (conv2d) layer and their parameters are optimized using a speech separation objective function in a purely data-driven fashion. Furthermore, inspired by the IPD formulation, we design a conv2d kernel to compute the inter-channel convolution differences (ICDs), which are expected to provide the spatial cues that help to distinguish the directional sources. Evaluation results on simulated multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed ICD based MCSS model improves the overall signal-to-distortion ratio by 10.4% over the IPD based MCSS model.

📄 PDF Abstract BibTeX arXiv:2003.03927

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Neural Speech Separation Using Spatially Distributed Microphones

2020-04-28 · Dongmei Wang, Zhuo Chen, Takuya Yoshioka

This paper proposes a neural network based speech separation method using spatially distributed microphones. Unlike with traditional microphone array settings, neither the number of microphones nor their spatial arrangem…

speech-recognitionSpeech RecognitionSpeech Separation

A comprehensive study of speech separation: spectrogram vs waveform separation

2019-05-17 · Fahimeh Bahmaninezhad, Jian Wu, Rongzhi Gu, Shi-Xiong Zhang 외

Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently, a raw audio waveform separation network…

speech-recognitionSpeech RecognitionSpeech Separation

Multi-channel Speech Separation Using Spatially Selective Deep Non-linear Filters

2023-04-24 · Kristina Tesch, Timo Gerkmann

In a multi-channel separation task with multiple speakers, we aim to recover all individual speech signals from the mixture. In contrast to single-channel approaches, which rely on the different spectro-temporal characte…

Speech Separation

Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure

2024-02-01 · Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri 외

We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our met…

Speech Enhancement

End-to-End Multi-Channel Speech Separation

2019-05-15 · Rongzhi Gu, Jian Wu, Shi-Xiong Zhang, Lian-Wu Chen 외

The end-to-end approach for single-channel speech separation has been studied recently and shown promising results. This paper extended the previous approach and proposed a new end-to-end model for multi-channel speech s…

Speech Separation