paper-with-me

Papers

Speech separation with large-scale self-supervised learning

2022-11-09 · Zhuo Chen, Naoyuki Kanda, Jian Wu, Yu Wu, Xiaofei Wang, Takuya Yoshioka, Jinyu Li, Sunit Sivasankaran, Sefik Emre Eskimez

Self-supervised learning (SSL) methods such as WavLM have shown promising speech separation (SS) results in small-scale simulation-based experiments. In this work, we extend the exploration of the SSL-based SS by massively scaling up both the pre-training data (more than 300K hours) and fine-tuning data (10K hours). We also investigate various techniques to efficiently integrate the pre-trained model with the SS network under a limited computation budget, including a low frame rate SSL model training setup and a fine-tuning scheme using only the part of the pre-trained model. Compared with a supervised baseline and the WavLM-based SS model using feature embeddings obtained with the previously released 94K hours trained WavLM, our proposed model obtains 15.9% and 11.2% of relative word error rate (WER) reductions, respectively, for a simulated far-field speech mixture test set. For conversation transcription on real meeting recordings using continuous speech separation, the proposed model achieves 6.8% and 10.6% of relative WER reductions over the purely supervised baseline on AMI and ICSI evaluation sets, respectively, while reducing the computational cost by 38%.

📄 PDF Abstract BibTeX arXiv:2211.05172

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeech Separation

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training

2020-10-29 · Sung-Feng Huang, Shun-Po Chuang, Da-Rong Liu, Yi-Chen Chen 외

Speech separation has been well developed, with the very successful permutation invariant training (PIT) approach, although the frequent label assignment switching happening during PIT training remains to be a problem wh…

Speaker SeparationSpeech EnhancementSpeech Separation

CSLNSpeech: solving extended speech separation problem with the help of Chinese sign language

2020-07-21 · Jiasong Wu, Xuan Li, Taotao Li, Fanman Meng 외

Previous audio-visual speech separation methods use the synchronization of the speaker's facial movement and speech in the video to supervise the speech separation in a self-supervised way. In this paper, we propose a mo…

Self-Supervised LearningSpeech Separation

Investigating self-supervised learning for speech enhancement and separation

2022-03-15 · Zili Huang, Shinji Watanabe, Shu-wen Yang, Paola Garcia 외

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a…

Self-Supervised LearningSpeech EnhancementSpeech Separation

Speech Separation with Pretrained Frontend to Minimize Domain Mismatch

2024-11-05 · Wupeng Wang, Zexu Pan, Xinke Li, Shuai Wang 외

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail pa…

Speech Separation

Probing Self-supervised Learning Models with Target Speech Extraction

2024-02-17 · Junyi Peng, Marc Delcroix, Tsubasa Ochiai, Oldrich Plchot 외

Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-talker scenarios, such as extracting a t…

Self-Supervised LearningSpeaker IdentificationSpeaker VerificationSpeech Extraction+1