paper-with-me

홈 › Papers

Single-Channel Multi-Speaker Separation using Deep Clustering

2016-07-07 · Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, John R. Hershey

Deep clustering is a recently introduced deep learning architecture that uses discriminatively trained embeddings as the basis for clustering. It was recently applied to spectrogram segmentation, resulting in impressive results on speaker-independent multi-speaker separation. In this paper we extend the baseline system with an end-to-end signal approximation objective that greatly improves performance on a challenging speech separation. We first significantly improve upon the baseline system performance by incorporating better regularization, larger temporal context, and a deeper architecture, culminating in an overall improvement in signal to distortion ratio (SDR) of 10.3 dB compared to the baseline of 6.0 dB for two-speaker separation, as well as a 7.1 dB SDR improvement for three-speaker separation. We then extend the model to incorporate an enhancement layer to refine the signal estimates, and perform end-to-end training through both the clustering and enhancement stages to maximize signal fidelity. We evaluate the results using automatic speech recognition. The new signal approximation objective, combined with end-to-end training, produces unprecedented performance, reducing the word error rate (WER) from 89.1% down to 30.8%. This represents a major advancement towards solving the cocktail party problem.

📄 PDF Abstract BibTeX arXiv:1607.02173

Code (2)

JusperLee/Deep-Clustering-for-Speech-Separation pytorch
ishandutta2007/Speech-Denoising-Landscape

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringDeep ClusteringSpeaker Separationspeech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Efficient Integration of Multi-channel Information for Speaker-independent Speech Separation

2020-08-11

Although deep-learning-based methods have markedly improved the performance of speech separation over the past few years, it remains an open question how to integrate multi-channel signals for speech separation. We propo…

Deep ClusteringOpen-Ended Question AnsweringSpeech SeparationTransfer Learning

DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions

2024-11-11 · Shu-Tong Niu, Jun Du, Ruo-Yu Wang, Gao-Bin Yang 외

We propose a single-channel Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework for back-end automatic speech recognition (ASR), combining neural speaker diarization (NSD) and speech separation (SS). Fir…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+4

Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings

2024-09-25 · Ruoyu Wang, Shutong Niu, Gaobin Yang, Jun Du 외

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustnes…

Clusteringspeaker-diarizationSpeaker DiarizationSpeech Separation

Joint Sound Source Separation and Speaker Recognition

2016-04-29 · Jeroen Zegers, Hugo Van hamme

Non-negative Matrix Factorization (NMF) has already been applied to learn speaker characterizations from single or non-simultaneous speech for speaker recognition applications. It is also known for its good performance i…

blind source separationSpeaker Recognition

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

2023-03-13 · Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, howeve…

Online ClusteringSpeaker SeparationSpeech Separation