paper-with-me

홈 › Papers

End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation

2019-10-30 · Yi Luo, Zhuo Chen, Nima Mesgarani, Takuya Yoshioka

An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to different indexing of the microphones with the same locations, while the latter requires the system to be able to process inputs with varying dimensions. Conventional optimization-based beamforming techniques satisfy these requirements by definition, while for deep learning-based end-to-end systems those constraints are not fully addressed. In this paper, we propose transform-average-concatenate (TAC), a simple design paradigm for channel permutation and number invariant multi-channel speech separation. Based on the filter-and-sum network (FaSNet), a recently proposed end-to-end time-domain beamforming system, we show how TAC significantly improves the separation performance across various numbers of microphones in noisy reverberant separation tasks with ad-hoc arrays. Moreover, we show that TAC also significantly improves the separation performance with fixed geometry array configuration, further proving the effectiveness of the proposed paradigm in the general problem of multi-microphone speech separation.

📄 PDF Abstract BibTeX arXiv:1910.14104

Code (2)

yluo42/TAC 공식 구현 pytorch
yoonsanghyu/FaSNet-TAC-PyTorch pytorch

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Location-based training for multi-channel talker-independent speaker separation

2021-10-08 · Hassan Taherian, Ke Tan, DeLiang Wang

Permutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we prop…

Speaker Separation

Low-Latency Speaker-Independent Continuous Speech Separation

2019-04-13 · Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao 외

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of whi…

speech-recognitionSpeech RecognitionSpeech Separation

Consistent ICA: Determined BSS meets spectrogram consistency

2020-05-20 · Kohei Yatabe

Multichannel audio blind source separation (BSS) in the determined situation (the number of microphones is equal to that of the sources), or determined BSS, is performed by multichannel linear filtering in the time-frequ…

blind source separation

Convolutive Audio Source Separation using Robust ICA and an intelligent evolving permutation ambiguity solution

2017-08-14 · Dimitrios Mallis, Thomas Sgouros, Nikolaos Mitianoudis

Audio source separation is the task of isolating sound sources that are active simultaneously in a room captured by a set of microphones. Convolutive audio source separation of equal number of sources and microphones has…

Audio Source Separation

Multi-channel Time-Varying Covariance Matrix Model for Late Reverberation Reduction

2019-10-19

In this paper, a multi-channel time-varying covariance matrix model for late reverberation reduction is proposed. Reflecting that variance of the late reverberation is time-varying and it depends on past speech source va…