paper-with-me

홈 › Papers

Wavesplit: End-to-End Speech Separation by Speaker Clustering

2020-02-20 · Neil Zeghidour, David Grangier

We introduce Wavesplit, an end-to-end source separation system. From a single mixture, the model infers a representation for each source and then estimates each source signal given the inferred representations. The model is trained to jointly perform both tasks from the raw waveform. Wavesplit infers a set of source representations via clustering, which addresses the fundamental permutation problem of separation. For speech separation, our sequence-wide speaker representations provide a more robust separation of long, challenging recordings compared to prior work. Wavesplit redefines the state-of-the-art on clean mixtures of 2 or 3 speakers (WSJ0-2/3mix), as well as in noisy and reverberated settings (WHAM/WHAMR). We also set a new benchmark on the recent LibriMix dataset. Finally, we show that Wavesplit is also applicable to other domains, by separating fetal and maternal heart rates from a single abdominal electrocardiogram.

📄 PDF Abstract BibTeX arXiv:2002.08933

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringData AugmentationSpeech Separation

Similar Papers 제목 키워드 기반

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

2023-03-13 · Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, howeve…

Online ClusteringSpeaker SeparationSpeech Separation

Separation Guided Speaker Diarization in Realistic Mismatched Conditions

2021-07-06 · Shu-Tong Niu, Jun Du, Lei Sun, Chin-Hui Lee

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) appro…

Clusteringspeaker-diarizationSpeaker DiarizationSpeech Separation

Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation

2024-03-27 · Xilin Jiang, Cong Han, Nima Mesgarani

Transformers have been the most successful architecture for various speech modeling tasks, including speech separation. However, the self-attention mechanism in transformers with quadratic complexity is inefficient in co…

MambaSpeech SeparationState Space Models

Online speaker diarization of meetings guided by speech separation

2024-01-30 · Elio Gruttadauria, Mathieu Fontaine, Slim Essid

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation mode…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1

Single-Channel Multi-Speaker Separation using Deep Clustering

2016-07-07 · Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe 외

Deep clustering is a recently introduced deep learning architecture that uses discriminatively trained embeddings as the basis for clustering. It was recently applied to spectrogram segmentation, resulting in impressive …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringDeep Clustering+4