paper-with-me

홈 › Papers

Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker Separation

2019-04-25 · Yuzhou Liu, DeLiang Wang

We address talker-independent monaural speaker separation from the perspectives of deep learning and computational auditory scene analysis (CASA). Specifically, we decompose the multi-speaker separation task into the stages of simultaneous grouping and sequential grouping. Simultaneous grouping is first performed in each time frame by separating the spectra of different speakers with a permutation-invariantly trained neural network. In the second stage, the frame-level separated spectra are sequentially grouped to different speakers by a clustering network. The proposed deep CASA approach optimizes frame-level separation and speaker tracking in turn, and produces excellent results for both objectives. Experimental results on the benchmark WSJ0-2mix database show that the new approach achieves the state-of-the-art results with a modest model size.

📄 PDF Abstract BibTeX arXiv:1904.11148

Code (1)

yuzhou-git/deep-casa tf

Tasks

ClusteringSpeaker SeparationSpeech Separation

Similar Papers 제목 키워드 기반

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks

2017-03-18 · Morten Kolbæk, Dong Yu, Zheng-Hua Tan, Jesper Jensen

In this paper we propose the utterance-level Permutation Invariant Training (uPIT) technique. uPIT is a practically applicable, end-to-end, deep learning based solution for speaker independent multi-talker speech separat…

ClusteringDeep ClusteringSpeech Separation

Divide and Conquer Local Average Regression

2016-01-23 · Xiangyu Chang, Shao-Bo Lin, Yao Wang

The divide and conquer strategy, which breaks a massive data set into a se- ries of manageable data blocks, and then combines the independent results of data blocks to obtain a final decision, has been recognized as a st…

regression

Divide-and-Conquer with Sequential Monte Carlo

2014-06-19 · Fredrik Lindsten, Adam M. Johansen, Christian A. Naesseth, Bonnie Kirkpatrick 외

We propose a novel class of Sequential Monte Carlo (SMC) algorithms, appropriate for inference in probabilistic graphical models. This class of algorithms adopts a divide-and-conquer approach based upon an auxiliary tree…

Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective

2018-11-22 · Zhong-Qiu Wang, Ke Tan, DeLiang Wang

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sou…

Speaker Separation